Task planning method and device for large language model, storage medium and equipment
By filtering and generating text summary, we improve the task planning method of the large language model, and combine tool description mapping relationships to make tool calls, solving the problems of low task planning efficiency and inaccurate tool use in the existing technology, achieving more efficient and accurate task execution and user interaction experience.
Patent Information
- Application Number
- CN202510243536.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-03
- Publication Date
- 2025-06-20
AI Technical Summary
The existing large language models have problems such as difficulty in splitting complex tasks and inaccurate tool use when planning tasks, resulting in low efficiency and accuracy in handling complex tasks and poor user interaction experience.
By obtaining the target task text, filtering out text blocks similar to it, generating a text summary, and inputting it into the large language model with prompt instructions to obtain a subtask list. Describe the mapping relationship based on the pre-built tool, call the tool that executes the subtask to obtain the final processing result.
It improves the accuracy of the large language model in task planning and tool call, improves the execution efficiency and quality of complex tasks, and improves the user interaction experience.
Smart Images

Figure CN120179816A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of natural language processing technology, and in particular, to a task planning method, device, storage medium, and equipment for large language models. Background Art
[0002] With the rapid development of information technologies such as artificial intelligence and the Internet of Things, the application scenarios of human-computer interaction are becoming more and more extensive. A variety of large language models (LLMs) appear in people's life and work, such as Chat Generative Pre-trained Transformer (ChatGPT), smart speakers, smart TVs, etc., which can provide intelligent interaction functions for many application scenarios such as task planning to assist users in completing various behavioral intentions.
[0003] Currently, when using the large language model LLM for task planning, the commonly used method is to first use the Retrieval-Augmented Generation (RAG) technology to retrieve relevant information from an external knowledge base to enhance the task planning and task-solving capabilities of the large language model, and then use the tool call method to enable the large language model to have the ability to solve various problems to achieve the effect of task planning. However, this method has problems such as difficulty in splitting complex tasks and inaccurate tool usage, which hinder the efficiency and accuracy of the large language model in processing complex tasks and also reduce the user's interaction experience. Summary of the Invention
[0004] The main purpose of the embodiments of this application is to provide a task planning method, device, storage medium, and equipment for large language models, which can effectively improve the task decomposition and planning capabilities of large language models, improve the accuracy of tool calls of large language models, and further improve the efficiency and correctness of large language model task planning and solving complex tasks, and enhance the user's interaction experience.
[0005] The embodiments of this application provide a task planning method for large language models, including:
[0006] Obtain the target task text to be planned; and determine the target field to which the target task belongs;
[0007] Select M text blocks similar to the target task text from the knowledge documents of the target field; M is a positive integer greater than 0;
[0008] For the text summarization generation method based on multi-dimensional prompts, the key information of the M text blocks is combined with the first prompt instruction "prompt" and input into a preset large language model to obtain the text summary of the M text blocks output by the large language model;
[0009] The text summary is combined with the second prompt instruction "prompt" and input into the large language model to obtain N subtasks corresponding to the target task output by the large language model; N is a positive integer greater than 0;
[0010] According to the pre-constructed tool description mapping relationship, the tools for executing the N subtasks are called through the large language model, and the processing results of the N subtasks are obtained by executing the tools to form the final processing result for the target task.
[0011] In a possible implementation manner, the screening of M text blocks similar to the target task text from the knowledge documents in the target field includes:
[0012] Using the retrieval window technology, the knowledge documents in the target field are finely divided, and M text blocks similar to the target task text are screened from the division results.
[0013] In a possible implementation manner, the using the retrieval window technology to finely divide the knowledge documents in the target field and screening M text blocks similar to the target task text from the division results includes:
[0014] The knowledge documents in the target field are finely divided into continuous text segments of a preset length;
[0015] By means of a sliding window, each continuous text segment is traversed to obtain the text block corresponding to each window, and the similarity score between the text block corresponding to each window and the target task text is calculated;
[0016] M text blocks with similarity scores higher than a preset threshold are screened out as M text blocks similar to the target task text.
[0017] In a possible implementation manner, the calculating the similarity score between the text block corresponding to each window and the target task text includes:
[0018] The similarity score between the text block corresponding to each window and the target task text is calculated through an embedding vector model.
[0019] In a possible implementation manner, the preset threshold is 0.8.
[0020] In a possible implementation, for the text summary generation method based on multi-dimensional cues, the key information of the M text blocks is combined with the first prompt instruction "prompt" and input into a preset large language model to obtain the text summary of the M text blocks output by the large language model, including:
[0021] The content of the M text blocks is screened from three dimensions of semantics, keywords, and themes, so that the key information of the M text blocks is combined with the first prompt instruction "prompt" and input into a preset large language model to obtain the text summary of the M text blocks output by the large language model.
[0022] In a possible implementation, the construction method of the tool description mapping relationship is as follows:
[0023] The tools are hierarchically managed and mapped through the tool call hierarchy mapping technology to obtain a tool description mapping tree; the tool description mapping tree is used to display the tool description mapping relationship.
[0024] In a possible implementation, according to the pre-constructed tool description mapping relationship, the large language model is used to call the tools for executing the N subtasks, and by executing the tools, the processing results of the N subtasks are obtained to constitute the final processing result for the target task, including:
[0025] When the large language model is used to call the tools corresponding to each step of the N subtasks, according to the tool description mapping relationship displayed by the tool description mapping tree, relevant context information is passed between different steps through hierarchical mapping to ensure that the call to the tool corresponding to each step is an effective parameter transfer based on the output of the previous step, so as to obtain the processing results of the N subtasks and constitute the final processing result for the target task.
[0026] In a possible implementation, when the large language model is used to call the tools corresponding to each step of the N subtasks, according to the tool description mapping relationship displayed by the tool description mapping tree, the relevant context information between different steps is passed to ensure that the call to the tool corresponding to each step is an effective parameter transfer based on the output of the previous step, so as to obtain the processing results of the N subtasks and constitute the final processing result for the target task, including:
[0027] The content of the i-th subtask among the N subtasks and the top-level description information of the tool description mapping tree are combined with the third prompt instruction "prompt" and input into the large language model to obtain the description information of the tool for executing the i-th subtask output by the large language model; i is a positive integer greater than 0 and not greater than N;
[0028] Combine the content of the \(i\)-th sub-task and the description information of the tool corresponding to the execution of the \(i\)-th sub-task, and input them into the large language model together with the fourth prompt instruction prompt, so as to call the tool corresponding to the \(i\)-th sub-task through the model to obtain the processing result of the \(i\)-th sub-task, and so on until the processing results of the \(N\) sub-tasks are obtained, which are used to constitute the final processing result of the target task.
[0029] The embodiment of the present application also provides a task planning device for a large language model, including:
[0030] An acquisition unit, configured to acquire the target task text to be planned; and determine the target field to which the target task belongs;
[0031] A screening unit, configured to screen out \(M\) text blocks similar to the target task text from the knowledge documents in the target field; \(M\) is a positive integer greater than 0;
[0032] A first input unit, configured to input the key information of the \(M\) text blocks into a preset large language model based on the text summary generation method of multi-dimensional prompts, in combination with the first prompt instruction prompt, to obtain the text summary of the large language model output for the \(M\) text blocks;
[0033] A second input unit, configured to input the text summary into the large language model in combination with the second prompt instruction prompt to obtain \(N\) sub-tasks corresponding to the target text output by the large language model; \(N\) is a positive integer greater than 0;
[0034] An obtaining unit, configured to call the tools for executing the \(N\) sub-tasks through the large language model according to the pre-constructed tool description mapping relationship, and obtain the processing results of the \(N\) sub-tasks by executing the tools, which are used to constitute the final processing result of the target task.
[0035] In a possible implementation manner, the screening unit is specifically configured to:
[0036] Use the retrieval window technology to perform fine-grained partitioning on the knowledge documents in the target field, and screen out \(M\) text blocks similar to the target task text from the partitioning results.
[0037] In a possible implementation manner, the screening unit includes:
[0038] A partitioning sub-unit, configured to perform fine-grained partitioning on the knowledge documents in the target field into continuous text segments of a preset length;
[0039] A calculation subunit, configured to calculate, by means of a sliding window, traverse each continuous text segment, obtain a text block corresponding to each window, and calculate a similarity score between the text block corresponding to each window and the target task text;
[0040] A screening subunit, configured to screen out M text blocks with similarity scores higher than a preset threshold as M text blocks similar to the target task text.
[0041] In a possible implementation manner, the calculation subunit is specifically configured to:
[0042] Calculate a similarity score between the text block corresponding to each window and the target task text through an embedding vector model.
[0043] In a possible implementation manner, the preset threshold is 0.8.
[0044] In a possible implementation manner, the first input unit is specifically configured to:
[0045] Screen the content of the M text blocks from three dimensions of semantics, keywords, and themes, so as to combine the key information of the M text blocks with the first prompt instruction prompt and input it into a preset large language model to obtain a text summary of the M text blocks output by the large language model.
[0046] In a possible implementation manner, the device further includes:
[0047] A mapping unit, configured to hierarchically manage and map tools through a tool call hierarchical mapping technology to obtain a tool description mapping tree; the tool description mapping tree is used to display tool description mapping relationships.
[0048] In a possible implementation manner, the obtaining unit is specifically configured to:
[0049] When calling and executing tools corresponding to each step in the N subtasks through a large language model, according to the tool description mapping relationships displayed by the tool description mapping tree, through hierarchical mapping, transfer relevant context information between different steps to ensure that the call to the tool corresponding to each step is an effective parameter transfer based on the output of the previous step, and then obtain the processing results of the N subtasks to constitute the final processing result for the target task.
[0050] In a possible implementation manner, the obtaining unit includes:
[0051] The first input subunit is configured to input the content of the i-th sub-task among the N sub-tasks and the top-level description information of the tool description tree, in combination with the third prompt instruction "prompt", into the large language model, so as to obtain the description information of the tool corresponding to the execution of the i-th sub-task output by the large language model; where i is a positive integer greater than 0 and not greater than N;
[0052] The second input subunit is configured to input the content of the i-th sub-task and the description information of the tool corresponding to the execution of the i-th sub-task, in combination with the fourth prompt instruction "prompt", into the large language model, so as to execute the tool corresponding to the i-th sub-task through model invocation to obtain the processing result of the i-th sub-task, and so on until the processing results of the N sub-tasks are obtained, which are used to constitute the final processing result of the target task.
[0053] An embodiment of the present application further provides a task planning device for a large language model, including: a processor, a memory, and a system bus;
[0054] The processor and the memory are connected through the system bus;
[0055] The memory is used to store one or more programs, and the one or more programs include instructions that, when executed by the processor, cause the processor to execute any implementation manner of the above-mentioned task planning method for the large language model.
[0056] An embodiment of the present application further provides a computer-readable storage medium, in which instructions are stored, and when the instructions are run on a terminal device, the terminal device is caused to execute any implementation manner of the above-mentioned task planning method for the large language model.
[0057] An embodiment of the present application further provides a computer program product, which, when run on a terminal device, causes the terminal device to execute any implementation manner of the above-mentioned task planning method for the large language model.
[0058] A task planning method, device, storage medium, and equipment for a large language model provided by an embodiment of the present application first obtain a target task text to be planned; determine the target domain to which the target task belongs; then, from the knowledge documents of the target domain, screen out M text blocks similar to the target task text; where M is a positive integer greater than 0; then, based on a text summarization generation method with multi-dimensional prompts, combine the key information of the M text blocks with a first prompt instruction prompt and input it into a preset large language model to obtain a text summary of the M text blocks output by the large language model, and then input the text summary into the large language model in combination with a second prompt instruction prompt to obtain N subtasks corresponding to the target text output by the large language model; where N is a positive integer greater than 0. Furthermore, according to a pre-constructed tool description mapping relationship, tools for executing the N subtasks can be called through the large language model, and by executing these tools, processing results of the N subtasks can be obtained to form a final processing result for the target task.
[0059] It can be seen that since the present application, when processing the target task, first effectively obtains external knowledge related to the target task text (i.e., the text summary of the M text blocks) through the method of screening out short text blocks from long knowledge documents and the text summarization ability with multi-dimensional prompts, and then according to the pre-constructed tool description mapping relationship, reduces the proportion of tool descriptions in the input context of the large language model, thereby improving the accuracy of tool calls of the large language model, further enhancing the overall efficiency and quality of the execution of the target task, and also enhancing the interaction experience of the user who proposes the target task to the large language model. BRIEF DESCRIPTION OF THE DRAWINGS
[0060] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0061] Figure 1 It is a schematic flowchart of a task planning method for a large language model provided by an embodiment of the present application;
[0062] Figure 2 It is a schematic diagram of the overall process of task planning for a large language model provided by an embodiment of the present application;
[0063] Figure 3 It is a schematic diagram of a tool description mapping tree provided by an embodiment of the present application;
[0064] Figure 4 It is a schematic diagram of the composition of a task planning device for a large language model provided by an embodiment of the present application. Detailed implementation manners
[0065] Task planning of large language models refers to their ability to automatically generate a series of reasonable steps and plans according to given tasks and goals to complete complex tasks. This requires the model to have various capabilities such as understanding of tasks and perception of the environment. With the continuous development of large language model technology, more and more work uses large language models to process complex tasks. However, relying solely on the capabilities of large language models themselves to directly solve complex problems often results in unreasonable and inaccurate answers. At the same time, with the development of large language model agents, task planning has become a key link in the agent process, directly affecting the steps, efficiency, and accuracy of agents in solving problems. Therefore, it is of great significance to study the task planning scheme of large language models, improve the task planning and tool invocation capabilities of large language models, optimize the rationality and correctness of large language models in solving complex tasks, and improve the task processing flow of large language models.
[0066] Currently, when using large language model LLM for task planning, one common approach is to first use retrieval-augmented generation technology (RAG) to retrieve relevant information from external knowledge bases to enhance the task planning and task-solving capabilities of large language models. However, even if the information retrieved by the large language model is relevant to the task, it may not be the most appropriate or most direct content to answer the question, such as the retrieval granularity or information mismatch. This information mismatch may result in redundant, irrelevant, or suboptimal information in the generated text, affecting the quality of the answer.
[0067] On the other hand, when using large language model LLM for task planning, the common method is to use tool invocation to enable the large language model to have the ability to solve various problems. Although this method can bring stronger functionality and task execution capabilities, this method requires relying on the understanding of task context. Once the number of tools increases, the proportion of tool descriptions and parameter descriptions in the context will continue to expand, which may lead to the large language model misinvoking tools or passing incorrect parameters; tool invocation and the language model are usually discrete, and when the model invokes tools, it may not be able to fully utilize the global context information of the task. For example, the model cannot effectively transfer relevant information of previous and subsequent task steps to the tool, resulting in low tool invocation efficiency or inaccurate results.
[0068] It can be seen that for simple tasks, only a small number of tools are needed, and the large language model can complete the task clearly and accurately with a high probability. However, as the difficulty of the task increases, the large language model faces complex task scenarios, the task granularity is not broken down finely, the task goal is not broken down accurately, and the tool acquisition is incorrect. It is very easy to cause the system execution steps to be unable to be exhausted, and the loop can never end, resulting in the large language model being unable to successfully complete the task. In addition, as the number of tools increases, due to the input length of the large language model, too many tools and parameter descriptions will make it difficult for the large language model to capture important context and details, thereby affecting the model's choice of tools and reducing the accuracy of the large language model in solving tasks.
[0069] Therefore, in order to improve the working efficiency of large language models and their ability to solve complex tasks, it is extremely important to decompose tasks in a refined manner and call tools accurately. How to improve the efficiency and accuracy of large language models in handling complex tasks, and thereby improve the user's interactive experience, is a technical problem that needs to be solved urgently.
[0070] To solve the above defects, the present application provides a task planning method for a large language model, firstly obtaining the target task text to be planned; and determining the target field to which the target task belongs; then screening out M text blocks similar to the target task text from the knowledge documents in the target field; wherein M is a positive integer greater than 0; then, based on a text summary generation method of multi-dimensional prompts, the key information of the M text blocks, combined with the first prompt instruction prompt, is input into a preset large language model to obtain a text summary of the M text blocks output by the large language model, and then the text summary, combined with the second prompt instruction prompt, is input into the large language model to obtain N subtasks corresponding to the target text output by the large language model; wherein N is a positive integer greater than 0. Then, according to the pre-built tool description mapping relationship, the tool for executing N subtasks can be called through the large language model, and by executing these tools, the processing results of the N subtasks are obtained to constitute the final processing result for the target task.
[0071] It can be seen that when processing the target task, the present application first effectively obtains external knowledge related to the target task text (i.e., the text summary of M text blocks) by screening out short text blocks from long knowledge documents and the text summarization capability of multi-dimensional prompts, and then reduces the proportion of tool description in the input context of the large language model according to the pre-built tool description mapping relationship, thereby improving the accuracy of the large language model tool call, thereby improving the overall efficiency and quality of the execution of the target task, and also improving the interactive experience of users who propose target tasks to the large language model.
[0072] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the following will clearly and completely describe the technical solutions in the embodiments of this application with reference to the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are part of the embodiments of this application, rather than all of them. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in this application without creative efforts belong to the scope of protection of this application.
[0073] First Embodiment
[0074] See Figure 1 , which is a schematic flowchart of a task planning method for a large language model provided in this embodiment. The method includes the following steps:
[0075] S101: Obtain the target task text to be planned; and determine the target field to which the target task belongs.
[0076] In this embodiment, any task text that needs to be planned input by the user into an intelligent interaction software or device such as a large language model (LLM) is defined as the target task text to be planned. Moreover, this embodiment does not limit the language type of the target task text. For example, the target task text can be a Chinese text or an English text, etc. This embodiment also does not limit the length of the target task text, that is, the target question text can be a sentence text (i.e., a set of words) or a passage text (i.e., a set of sentences), etc. This embodiment also does not limit the specific content and type of the target task, which can be set according to actual situations and empirical values. For example, the target task can include, but is not limited to, word- or sentence-based association tasks, context-based automatic dialogue tasks, etc. For example, the target task text can be the question text "How many people are there in the capital city of the country that produces Model A engines" proposed by the user to the large language model.
[0077] Among them, the large language model LLM can be a language model based on deep learning, which can generate new language expressions according to the input text content, such as text, sentences, paragraphs, or even articles, etc. The large language model LLM is trained using a large-scale language dataset through an autoregressive generation method to obtain language rules and patterns, and can simulate human instructions to generate language expressions (such as text data). Specifically, when the large language model LLM generates new text data, it predicts the probability of the next language unit based on the content that has been generated previously until the complete text data is generated.
[0078] Further, after obtaining the target task text to be planned input by the user, existing or future text recognition algorithms can be used to perform domain recognition on the target task text to determine the domain to which the target task belongs (such as the education domain, the medical domain, the sports domain, etc.), and define it as the target domain for subsequent step S102.
[0079] S102: From the knowledge documents in the target domain, screen out M text blocks similar to the target task text; where M is a positive integer greater than 0.
[0080] In this embodiment, after obtaining the target task text to be planned through step S101 and determining the target domain to which the target task belongs, in order to improve the efficiency and correctness of the large language model in planning and solving the target task, and thus enhance the interaction experience of the user who proposes the target task, the present application proposes to first perform a fine-grained division of the knowledge documents in the target domain by screening out short text blocks from long knowledge documents (such as the retrieval window technique), obtaining a division result (including multiple short text blocks), and using existing or future text similarity calculation methods to calculate the similarity between the multiple short text blocks included in the division result and the target task text as the screening basis to screen out M text blocks more similar to the target task text from the division result. Then, by further executing subsequent steps and utilizing the text summarization ability of multi-dimensional prompts to effectively obtain external knowledge related to the target task text (i.e., the text summaries corresponding to the subsequent-mentioned M text blocks), the proportion of the tool description in the input context of the large language model can be reduced according to the pre-constructed tool description mapping relationship, thereby improving the accuracy of tool invocation of the large language model, and further enhancing the overall efficiency and quality of target task execution, as well as the interaction experience of the user who proposes the target task to the large language model, as Figure 2 shown.
[0081] It should be noted that since the content retrieved using the RAG technique in existing methods is not always exactly matched with the generation task. Therefore, in order to improve the relevance between the retrieved content and the target task, after obtaining the target task text to be planned and determining the target domain to which the target task belongs, the present application divides the knowledge documents in the target domain into multiple small pieces, and then combines the retrieval window technique (such as the sliding window mechanism) and relevance scoring to ensure that the retrieval system can accurately and quickly extract information related to the target task text when processing large-scale text for subsequent step S103.
[0082] Specifically, an optional implementation method is that after obtaining the target task text to be planned and determining the target domain to which the target task belongs, further, first, the knowledge documents in the target domain can be finely divided into several consecutive text segments of a preset length (fixed size or dynamically adjustable size), such as several consecutive text segments (which can also be understood as windows) obtained by dividing based on sentences, paragraphs, or a certain number of characters (such as 200 characters) / words.
[0083] Then, by means of a sliding window, each consecutive text segment can be traversed to obtain the text block corresponding to each window, and the similarity score between the text block corresponding to each window and the target task text can be calculated by means of an embedding vector model or the like.
[0084] Among them, by means of a sliding window (such as the window length can be 10 sentences, and the step length can be 8 sentences, etc.), that is, there is a certain overlap between each window. The sliding window mechanism can avoid breaks between consecutive text segments, thus ensuring that context information is better retained during the retrieval process. On this basis, the similarity score between the text block in each window and the target task text can be calculated. For example, through natural language processing technology or an embedding vector model, using the method of calculating text embedding vector similarity, the similarity score between the text block (which can be converted into a text embedding vector) in each window and the target task text (which can be converted into a text embedding vector) can be calculated, and scores such as 0.9 or 0.6 can be obtained.
[0085] Next, M (where M is a positive integer greater than 0) text blocks with similarity scores higher than a preset threshold (the specific value is not limited and can be set according to actual situations and empirical values, such as the preset threshold can be set to 0.8, etc.) can be selected as the M text blocks similar to the target task text.
[0086] In this way, by dividing the long knowledge documents in the target domain into multiple small pieces, combining the sliding window mechanism and relevance scoring, the relevance between the retrieved content and the task is effectively improved, so that when the large language model executes the target task later, the context information provided by these M text blocks can be utilized to ensure that the execution result can be as accurate as possible and obtained based on the actual retrieved content.
[0087] S103: Based on the text summary generation method with multi-dimensional prompts, the key information of the M text blocks is combined with the first prompt instruction prompt and input into a preset large language model to obtain the text summary of the M text blocks output by the large language model.
[0088] It should be noted that considering the context length limit of the input of the large language model in this application, providing all M text block contents similar to the target task text screened out through step S102 to the context of the large language model may exceed the context window limit of the model, resulting in some information being unable to be effectively processed. In addition, inputting too much content may increase the cognitive burden of the model and weaken its ability to understand the key points of the task. In such a case, the model is prone to falling into information overload, unable to focus on the most critical part, thereby affecting the quality and accuracy of the generated content.
[0089] Therefore, in this embodiment, the contents of the M text blocks related to the target task text screened out are screened, compressed, and refined from three dimensions: semantics, keywords, and themes, so as to retain the most important information. This not only ensures that the model can obtain key information within the limited context length but also effectively improves the model's understanding ability of the task. Then, by integrating the key information of the M text blocks into the first prompt instruction prompt and inputting it into the preset large language model, a compact and accurate text summary of the M text blocks output by the large language model can be obtained for subsequent execution of step S104.
[0090] Among them, screening from the semantic dimension is to focus on the overall meaning and logical structure of the corresponding text block and extract key information by understanding the context and core viewpoints of the text block. The specific screening method adopted is not limited and can be selected according to the actual situation and empirical values. For example, the entire text block can be first converted into an overall semantic vector to capture the overall meaning of the text block, and then each sentence in the text block is respectively converted into a corresponding semantic vector to calculate the similarity between the semantic information corresponding to each sentence and the overall semantic vector of the text block. The sentences with high similarity are screened out, and then, according to factors such as the semantic similarity and position of the screened sentences (for example, the sentences at the beginning and end are usually more important), an importance score is given to each sentence, and the sentences with higher scores are selected as one of the important bases for subsequent generation of the text summary.
[0091] Filtering from the keyword dimension is to focus on the words with high frequencies and close relevance to the theme in the corresponding text blocks. Through keyword extraction, the important information of the text blocks can be quickly located to ensure that the subsequent generated abstracts can contain core terms and improve the information concentration. The specific filtering methods are not limited and can be selected according to the actual situation and empirical values. For example, the keyword extraction function of the LLM can be used first, or keywords in the corresponding text blocks can be extracted based on algorithms such as Term Frequency-Inverse Document Frequency (TF-IDF) and TextRank, and then sentences containing these keywords are filtered out from the text blocks to ensure that the subsequent generated abstracts can cover the core content of the text blocks. The keyword coverage rate (such as the number / weight of keywords included) in each of the filtered sentences can be calculated to further select sentences with higher keyword coverage rates as one of the important bases for subsequent text abstract generation.
[0092] Moreover, filtering from the theme dimension is to focus on the core themes and main ideas of the corresponding text blocks, and extract sentences related to the theme by identifying the theme to avoid information bias. The specific filtering methods are still not limited and can be selected according to the actual situation and empirical values. For example, the theme annotation function of the LLM can be used first, or themes in the corresponding text blocks can be extracted based on algorithms such as LDA and Zero-shot Classification, and then sentences highly relevant to the theme are filtered out from the text blocks to ensure that each theme has enough sentences extracted as one of the important bases for subsequent text abstract generation to ensure the comprehensiveness of the subsequent generated abstracts.
[0093] On this basis, the screening results of the key information of M text blocks obtained from the three dimensions of semantics, keywords, and themes can be further integrated, and the text abstracts corresponding to each of the M text blocks can be efficiently generated through the large language model. For example, according to the requirements of the target task, weights can be assigned to the sentences filtered from the three dimensions of semantics, keywords, and themes. For example, weights of 0.5, 0.3, and 0.2 can be assigned to the sentences corresponding to the three dimensions respectively, and then the sentences are scored using these weights. The large language model is used to process the sentences with higher scores for redundancy removal, coherent recombination, etc., to generate the text abstracts corresponding to the M text blocks, so as to not only capture the core content of the text blocks but also ensure the accuracy, coherence, and comprehensiveness of the finally generated text abstracts.
[0094] It should be noted that this application does not limit the screening methods used for screening from the three dimensions of semantics, keywords, and themes, nor does it limit the execution order of screening from these three dimensions. For example, screening can be first performed from the "semantics" dimension, and then sequentially from the "keywords" and "themes" dimensions to gradually extract more accurate key information. Or screening can also be performed simultaneously from the three dimensions, and then the screening results can be integrated for subsequent processing, etc. In practical applications, the screening method can be selected and the screening order can be set according to the actual situation and experience. The specific implementation process will not be elaborated here one by one.
[0095] Specifically, an optional implementation method is that in order to reduce the length occupied by the retrieved content in the large language model's prompt instruction while maintaining the important information in the retrieved information, first, semantic analysis can be performed on the content of all M screened text blocks to determine their relevance to the target task objective. For highly relevant content, clustering processing is then performed to merge similar content into a single semantic unit to reduce duplicate and redundant information. Then, key information, including keywords and topic content, is further extracted from each cluster or text. The extracted key content and the screened text will be used as the main part of generating the abstract to replace the original lengthy text, forming an abstract generation prompt word (as Figure 2 shown, that is, the first prompt instruction prompt). This prompt word can cover the key information required for the task and avoid information redundancy, thereby ensuring that the large language model can handle the core task well within the limited context window.
[0096] In this way, by analyzing and processing the content of the M text blocks from multiple dimensions such as semantics, keywords, and themes, a compact and accurate text abstract is generated, and these text abstracts are used as prompts to be passed to the large language model. It can cover the key information required to execute the target task and avoid information redundancy, ensuring that the large language model can better handle the core task within the limited context window.
[0097] S104: Input the text abstract, combined with the second prompt instruction prompt, into the large language model to obtain N subtasks corresponding to the target text output by the large language model; where N is a positive integer greater than 0.
[0098] In this embodiment, after obtaining the text abstracts of the M text blocks through step S103, the text abstract and the target task text can be further incorporated into the task planning prompt word (as Figure 2 shown, which is defined here as the second prompt instruction prompt) and input into the large language model, and then N subtasks corresponding to the decomposition of the target task output by the large language model can be obtained for executing the subsequent step S105. Where N takes a positive integer greater than 0.
[0099] For example: Suppose the target task text is the question text "How many people are there in the capital of the country that produces Engine Model A" asked by the user to the large language model. At this time, the first prompt can be written in many ways according to the content of M text blocks. The examples can be as follows:
[0100] "# Given text
[0101] The Model A series of aircraft is a medium- and short-range twin-engine jet airliner produced by Company B in Country A. Since the design was launched in xx, it has gone through multiple development stages, including the traditional type, the improved type, and the latest XX series. The evolution of Engine Model A reflects the continuous progress of aviation industry technology and the increasing requirements for efficiency, safety, and environmental protection. ...
[0103] # Task
[0104] Given a piece of text, generate its summary. Ensure the clarity of the semantics, keywords, and themes of the summary content.".
[0105] After inputting this first prompt into the large language model LLM, an example of the text summary output by the model can be: "The Model A series of aircraft is a medium- and short-range twin-engine jet airliner produced by Company B in Country A. Since the design was launched in xx, it has gone through multiple development stages, including the traditional type, the improved type, and the latest XX series.". Further, based on this text summary and the target task text, an example of the task planning prompt word (i.e., the second prompt) can be as follows:
[0106] "# User's question
[0107] How many people are there in the capital of the country that produces Engine Model A
[0108] # Text summary
[0109] The Model A series of aircraft is a medium- and short-range twin-engine jet airliner produced by Company B in Country A. Since the design was launched in xx, it has gone through multiple development stages, including the traditional type, the improved type, and the latest XX series....
[0110] # Task:
[0111] Based on the context, write a plan or modify an existing plan to achieve the goal.
[0112] - The plan consists of 1 to 3 tasks. If you are modifying an existing plan, please follow the instructions carefully and do not make unnecessary modifications.
[0113] - Unless instructed to only modify a certain task in the plan, please provide the entire plan.
[0114] - If an error is encountered in the current task, simply correct it and output the current single task.
[0115] The output result should be a JSON list in the following format:
[0116]
[0117] After inputting this second prompt instruction "prompt" into the large language model LLM, the subtask examples output by the model can be as follows:
[0118]
[0119] S105: According to the pre-constructed tool description mapping relationship, call the tools for executing N subtasks through the large language model, and obtain the processing results of the N subtasks by executing the tools to form the final processing result for the target task.
[0120] It should be noted that in the process of the large language model solving tasks, tool invocation plays a key role. Especially when dealing with complex and diverse tasks, the effective invocation of external tools can significantly improve the model's processing ability and task execution efficiency. However, with the increase in the number of available tools and the complexity of their parameter descriptions, the amount of information that the model needs to process in the context also increases significantly. This makes the proportion of tools and parameter descriptions in the model input context relatively high. Therefore, in order to optimize the high proportion of tools and parameter descriptions in the context of the large language model input during the tool invocation process, which affects the model's understanding of tasks and judgment of tool invocation, this application proposes to hierarchically manage and map tools through the tool invocation hierarchy mapping technology to obtain a tool description mapping tree, as Figure 3 shown, to display the tool description mapping relationship, so as to help the model more clearly understand the relationship between tools and their invocation order.
[0121] Specifically, in the process of constructing the tool description mapping tree in this application, first, the existing tools are classified according to their functions, application scenarios, or task types, and divided into different levels according to the tree structure. Each level represents a different type of tool, ensuring that the model can effectively select the appropriate tool according to the context of the task when invoking the tool. Secondly, the definition and parameter description of each tool need to be clear and accurate, and need to be mapped to the corresponding tool parameter description through the tool name.
[0122] On this basis, after obtaining N subtasks corresponding to the target task through step S104, an optional implementation method is that when calling and executing the tools corresponding to each step in the N subtasks through a large language model, according to the tool description mapping relationship shown in the tool description mapping tree, through hierarchical mapping, relevant context information can be passed between different steps to ensure that the call to the tool corresponding to each step can perform effective parameter passing based on the output of the previous step. This not only improves the accuracy of tool calls but also maintains the consistency of the context during task execution, thereby obtaining more accurate processing results for the N subtasks, which are used to form the final processing result for the target task.
[0123] Specifically, first, the content of the i-th (where i is a positive integer greater than 0 and not greater than N, such as Figure 2 the tool call situation where i takes the value of 1 as shown) subtask among the N subtasks and the top-level description information of the tool description mapping tree can be incorporated into the third prompt instruction prompt (such as Figure 2 the "tool description prompt word" as shown), and input into the large language model to obtain the description information of the tool corresponding to executing the i-th subtask output by the large language model. Then, the content of the i-th subtask and the description information of the tool corresponding to executing the i-th subtask can be incorporated into the fourth prompt instruction prompt (such as Figure 2 the "tool call prompt word" as shown), and input into the large language model to call and execute the tool corresponding to the i-th subtask through the model to obtain the processing result of the i-th subtask, and so on, until the processing results of the N subtasks are obtained, which are used to form the final processing result for the target task.
[0124] In this implementation method, to solve the target task proposed by the user, it is necessary to solve each of the N subtasks. Therefore, it is necessary to query tools for each subtask accordingly. Taking subtask 1 (an information query type task) as an example, first, the subtask content and the top-level description information of the tool description mapping tree can be concatenated to the tool description prompt word, and the large language model can be used to judge the type of tool description to be called, such as an information query type tool description. Secondly, the information query type tool description and the subtask content can be given to the tool call prompt word and the large language model can be used to select the tool and parameters to be called. Finally, the tool is executed to obtain the result of subtask 1 and stored in the context information for the processing of the next subtask, and so on, and the processing results of all N subtasks can be obtained, which are used to form the final processing result for the target task.
[0125] In this way, by introducing the tool call hierarchy mapping technology, the tools are hierarchically managed, helping the model clearly understand the relationships between tools and ensuring that it can select the most appropriate tool when performing tasks. At the same time, through the transmission of context information, the model can flexibly call tools according to the context information of the task, reducing the possibility of redundant or incorrect tool calls. As a result, the large language model can achieve more efficient and accurate tool scheduling when processing target tasks, improving the overall task execution efficiency.
[0126] In summary, for the task planning method of a large language model provided in this embodiment, first, obtain the target task text to be planned; and determine the target domain to which the target task belongs; then, from the knowledge documents of the target domain, screen out M text blocks similar to the target task text; where M is a positive integer greater than 0; next, based on the text summarization method with multi-dimensional prompts, input the key information of the M text blocks, combined with the first prompt instruction prompt, into the preset large language model to obtain the text summary of the M text blocks output by the large language model, and then input this text summary, combined with the second prompt instruction prompt, into the large language model to obtain N subtasks corresponding to the target text output by the large language model; where N is a positive integer greater than 0. Furthermore, according to the pre-constructed tool description mapping relationship, the tools for executing the N subtasks can be called through the large language model, and by executing these tools, the processing results of the N subtasks can be obtained to constitute the final processing result of the target task.
[0127] It can be seen that when processing the target task in this application, first, by screening out short text blocks from long knowledge documents and the text summarization ability with multi-dimensional prompts, external knowledge related to the target task text is effectively obtained (i.e., the text summary of the M text blocks), and then according to the pre-constructed tool description mapping relationship, the proportion of tool descriptions in the input context of the large language model is reduced, thereby improving the accuracy of tool calls of the large language model, and further improving the overall efficiency and quality of target task execution, as well as the interaction experience of users who submit target tasks to the large language model.
[0128] Second Embodiment
[0129] This embodiment will introduce a task planning device for a large language model. For related content, please refer to the above method embodiment.
[0130] See Figure 4 , which is a schematic diagram of the composition of a task planning device for a large language model provided in this embodiment. The device 400 includes:
[0131] An acquisition unit 401, configured to obtain the target task text to be planned; and determine the target domain to which the target task belongs;
[0132] A screening unit 402, configured to screen out M text blocks similar to the target task text from the knowledge documents in the target field; M is a positive integer greater than 0;
[0133] A first input unit 403, configured to input the key information of the M text blocks, combined with a first prompt instruction prompt, into a preset large language model based on a text summarization generation method with multi-dimensional prompts, to obtain a text summary of the M text blocks output by the large language model;
[0134] A second input unit 404, configured to input the text summary, combined with a second prompt instruction prompt, into the large language model, to obtain N subtasks corresponding to the target text output by the large language model; N is a positive integer greater than 0;
[0135] An obtaining unit 405, configured to, according to a pre-constructed tool description mapping relationship, call a tool for executing the N subtasks through the large language model, and obtain processing results of the N subtasks by executing the tool, so as to constitute a final processing result for the target task.
[0136] In an implementation manner of this embodiment, the screening unit 402 is specifically configured to:
[0137] Use the retrieval window technology to perform fine-grained partitioning on the knowledge documents in the target field, and screen out M text blocks similar to the target task text from the partitioning results.
[0138] In an implementation manner of this embodiment, the screening unit 402 includes:
[0139] A partitioning subunit, configured to perform fine-grained partitioning on the knowledge documents in the target field into continuous text segments of a preset length;
[0140] A calculating subunit, configured to calculate, by means of a sliding window, each continuous text segment to obtain a text block corresponding to each window, and calculate a similarity score between the text block corresponding to each window and the target task text;
[0141] A screening subunit, configured to screen out M text blocks with similarity scores higher than a preset threshold as M text blocks similar to the target task text.
[0142] In an implementation manner of this embodiment, the calculating subunit is specifically configured to:
[0143] Calculate the similarity score between the text block corresponding to each window and the target task text through an embedding vector model.
[0144] In one implementation of this embodiment, the preset threshold is 0.8.
[0145] In one implementation of this embodiment, the first input unit 403 is specifically configured to:
[0146] Screen the content of the M text blocks from three dimensions of semantics, keywords, and themes, so as to combine the key information of the M text blocks with the first prompt instruction prompt and input it into a preset large language model to obtain a text summary of the M text blocks output by the large language model.
[0147] In one implementation of this embodiment, the device further includes:
[0148] A mapping unit, configured to hierarchically manage and map tools through tool call hierarchical mapping technology to obtain a tool description mapping tree; the tool description mapping tree is used to display tool description mapping relationships.
[0149] In one implementation of this embodiment, the obtaining unit 405 is specifically configured to:
[0150] When calling and executing the tools corresponding to the steps in the N subtasks through the large language model, according to the tool description mapping relationships shown in the tool description mapping tree, through hierarchical mapping, transfer relevant context information between different steps to ensure that the call of the tool corresponding to each step is an effective parameter transfer based on the output of the previous step, and then obtain the processing results of the N subtasks, which are used to constitute the final processing result of the target task.
[0151] In one implementation of this embodiment, the obtaining unit 405 includes:
[0152] A first input subunit, configured to combine the content of the i-th subtask in the N subtasks with the top-level description information of the tool description mapping tree and input it into the large language model together with the third prompt instruction prompt to obtain the description information of the tool corresponding to executing the i-th subtask output by the large language model; i is a positive integer greater than 0 and not greater than N;
[0153] A second input subunit, configured to combine the content of the i-th subtask with the description information of the tool corresponding to executing the i-th subtask and input it into the large language model together with the fourth prompt instruction prompt to execute the tool corresponding to the i-th subtask through model call to obtain the processing result of the i-th subtask, and so on until the processing results of the N subtasks are obtained, which are used to constitute the final processing result of the target task.
[0154] Furthermore, an embodiment of the present application further provides a task planning device for a large language model, including: a processor, a memory, and a system bus;
[0155] The processor and the memory are connected through the system bus;
[0156] The memory is used to store one or more programs, and the one or more programs include instructions that, when executed by the processor, cause the processor to execute any implementation method of the above-mentioned task planning method for the large language model.
[0157] Furthermore, an embodiment of the present application further provides a computer-readable storage medium, in which instructions are stored, and when the instructions are run on a terminal device, the terminal device is caused to execute any implementation method of the above-mentioned task planning method for the large language model.
[0158] Furthermore, an embodiment of the present application further provides a computer program product, which, when running on a terminal device, causes the terminal device to execute any implementation method of the above-mentioned task planning method for the large language model.
[0159] From the description of the above embodiments, those skilled in the art can clearly understand that all or part of the steps in the above embodiment methods can be implemented by means of software plus a necessary general hardware platform. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network communication device such as a media gateway, etc.) to execute the methods described in various embodiments or some parts of the embodiments of the present application.
[0160] It should be noted that the various embodiments in this specification are described in a progressive manner, and the key point of each embodiment is to describe the differences from other embodiments. The same or similar parts among the various embodiments can be referred to each other. For the devices disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple, and the relevant parts can be referred to the description of the method part.
[0161] It should also be noted that in this text, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or device comprising the element.
[0162] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present application. Various modifications to these embodiments will be apparent to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application will not be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A task planning method for a large language model, characterized in that: include: Get the target task text to be planned; and determine the target area to which the said target task belongs; Filter out M text blocks similar to the target task text from the knowledge documents in the target domain; M is a positive integer greater than 0; Based on the multi-dimensional prompt text summary generation method, the key information of the M text blocks is combined with the first prompt instruction prompt, and input into a preset large language model to obtain the text summary for the M text blocks output by the large language model; Input the text summary, combined with the second prompt instruction prompt, into the large language model to obtain N subtasks corresponding to the target task output by the large language model; N is a positive integer greater than 0; According to the pre-built tool description mapping relationship, the tool for executing the N subtasks is called through the large language model, and the processing results of the N subtasks are obtained by executing the tool to constitute the final processing result for the target task.
2. The method according to claim 1, characterized in that The step of selecting M text blocks similar to the target task text from the knowledge documents in the target domain includes: The retrieval window technology is used to perform fine-grained division of the knowledge documents in the target domain, and M text blocks similar to the target task text are screened out from the division results.
3. The method according to claim 2, characterized in that The retrieval window technology is used to perform fine-grained division of the knowledge documents in the target field, and M text blocks similar to the target task text are screened out from the division results, including: Fine-grained division of the knowledge documents in the target domain into continuous text segments of preset lengths; By sliding the window, traverse each continuous text segment to obtain the text block corresponding to each window, and calculate the similarity score between the text block corresponding to each window and the target task text; M text blocks with similarity scores higher than a preset threshold are screened out as the M text blocks similar to the target task text.
4. The method according to claim 1, characterized in that: The text summary generation method based on multi-dimensional prompts inputs the key information of the M text blocks into a preset large language model in combination with a first prompt instruction prompt, and obtains a text summary for the M text blocks output by the large language model, including: The contents of the M text blocks are screened from three dimensions: semantics, keywords, and themes, so that the key information of the M text blocks is input into a preset large language model in combination with a first prompt instruction prompt, and a text summary for the M text blocks output by the large language model is obtained.
5. The method according to any one of claims 1 to 4, characterized in that: The tool describes how the mapping relationship is constructed as follows: The tools are hierarchically managed and mapped by using the tool call hierarchical mapping technology to obtain a tool description mapping tree; the tool description mapping tree is used to display the tool description mapping relationship.
6. The method according to claim 5, characterized in that The tool describing the mapping relationship according to the pre-built tool is called through the large language model to execute the N subtasks, and the processing results of the N subtasks are obtained by executing the tool to form the final processing result for the target task, including: When the tools corresponding to each step in the N subtasks are called through the large language model, the relevant context information is passed between different steps through hierarchical mapping based on the tool description mapping relationship displayed in the tool description mapping tree to ensure that the call to the tool corresponding to each step is a valid parameter transfer based on the output of the previous step, thereby obtaining the processing results of the N subtasks to constitute the final processing result of the target task.
7. The method according to claim 6, characterized in that When the tool corresponding to each step in the N subtasks is called and executed through the large language model, the relevant context information between different steps is transferred according to the tool description mapping relationship displayed by the tool description mapping tree to ensure that the call of the tool corresponding to each step is a valid parameter transfer based on the output of the previous step, thereby obtaining the processing results of the N subtasks to constitute the final processing result of the target task, including: Input the content of the i-th subtask among the N subtasks and the top-level description information of the tool description mapping tree into the large language model in combination with the third prompt instruction prompt, and obtain the description information of the tool corresponding to the i-th subtask output by the large language model; i is a positive integer greater than 0 and not greater than N; The content of the i-th subtask and the description information of the tool corresponding to the i-th subtask are input into the large language model in combination with the fourth prompt instruction prompt, so as to execute the tool corresponding to the i-th subtask through model call, so as to obtain the processing result of the i-th subtask, and so on, until the processing results of the N subtasks are obtained, so as to constitute the final processing result for the target task.
8. A task planning device for a large language model, characterized in that: include: An acquisition unit, used to acquire the target task text to be planned; and determine the target area to which the said target task belongs; A screening unit, used to screen out M text blocks similar to the target task text from the knowledge documents in the target domain; M is a positive integer greater than 0; A first input unit is used for inputting the key information of the M text blocks into a preset large language model in combination with a first prompt instruction prompt based on a text summary generation method of a multi-dimensional prompt, so as to obtain a text summary for the M text blocks output by the large language model; A second input unit is used to input the text summary, combined with a second prompt instruction prompt, into the large language model to obtain N subtasks corresponding to the target text output by the large language model; N is a positive integer greater than 0; The obtaining unit is used to describe the mapping relationship according to the pre-built tool, call the tool for executing the N subtasks through the large language model, and obtain the processing results of the N subtasks by executing the tool to constitute the final processing result for the target task.
9. A task planning device for a large language model, characterized in that: include: Processor, memory, system bus; The processor and the memory are connected via the system bus; The memory is used to store one or more programs, wherein the one or more programs include instructions, and when the instructions are executed by the processor, the processor executes the method according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores instructions, and when the instructions are executed on a terminal device, the terminal device executes the method according to any one of claims 1 to 7.
Citation Information
Cited By
Model prompt content generation method and device
CN120611029A
A model prompt content generation method and device
CN120611029B
Task execution method, device and system, electronic device and storage medium
CN120723411A
Task execution method based on intelligent agent
CN120892475A
An agent-based task execution method
CN120892475B