Prompt information determination method and device, intelligent agent, electronic equipment, storage medium and program product
By splitting the file into content blocks and determining the correlation, the target element block construction prompt information is selected, which solves the problem of information loss when the input limit is exceeded, and efficient and complete reply generation is achieved.
Patent Information
- Application Number
- CN202510804002.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-16
- Publication Date
- 2025-09-02
AI Technical Summary
When large models process file content that exceeds the maximum input limit, it is difficult to efficiently extract file content related to instructions, resulting in the loss of key information or generate incomplete replies.
By splitting the target file into content blocks and retaining the structure layout information, the correlation between the task description information and element blocks is determined, the target element blocks are filtered out, and the prompt information is constructed to guide the big model to generate reply content.
Maintain semantic coherence and task adaptability during the compression process, improve the integrity of key content and structural layout information, and ensure the accuracy and completeness of the reply content generated by the big model.
Smart Images

Figure CN120579644A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, particularly large language models, intelligent agents, and human-computer interaction. It can be applied to scenarios such as file processing, document reading, intelligent customer service, and question-and-answering. More specifically, this application provides a method, apparatus, intelligent agent, electronic device, storage medium, and program product for determining prompt information. Background Art
[0002] With the continuous development of artificial intelligence technology, large models have made significant progress in natural language processing in recent years. By inputting user commands (queries) and file content, large models can generate corresponding responses. However, there are limits on the number of tokens (i.e., the size of the context window) that large models can process. Within this limit, the relevance of the input file content to the command directly affects the large model's response capabilities and applicable scenarios. Summary of the Invention
[0003] The present application provides a method, device, intelligent agent, electronic device, storage medium and program product for determining prompt information.
[0004] According to one aspect of the present application, a method for determining prompt information is provided, comprising: determining multiple correlations between task description information of a target object and multiple element blocks of a target file, the element blocks comprising content blocks split from the target file according to the file structure layout and structural layout information associated with the content blocks; based on the multiple correlations, screening out at least one target element block from the multiple element blocks; based on the task description information and the at least one target element block, determining prompt information for a large model, the prompt information being used to guide the large model to generate reply content for the task description information.
[0005] According to another aspect of the present application, a prompt information determination device is provided, including: a correlation module for determining multiple correlations between task description information of a target object and multiple element blocks of a target file, the element blocks including content blocks split from the target file according to the file structure layout and structural layout information associated with the content blocks; a screening module for screening out at least one target element block from the multiple element blocks based on multiple correlations; a prompt information module for determining prompt information for a large model based on the task description information and the at least one target element block, the prompt information being used to guide the large model to generate reply content for the task description information.
[0006] According to another aspect of the present application, an intelligent agent is provided, including an input module for receiving input information, the input information including prompt information obtained by executing the prompt information determination method provided in the embodiment of the present application; a processing module for obtaining output information by calling a large model based on the input information received by the input module; and an output module for outputting the output information obtained by the processing module.
[0007] According to another aspect of the present application, an electronic device is provided, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so as to enable the at least one processor to execute the prompt information determination method provided in accordance with an embodiment of the present application.
[0008] According to another aspect of the present application, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to enable a computer to execute the prompt information determination method provided in the present application.
[0009] According to another aspect of the present application, a computer program product is provided, including a computer program, which implements the prompt information determination method provided in the present application when executed by a processor.
[0010] It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present application, nor is it intended to limit the scope of the present application. Other features of the present application will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] The accompanying drawings are used to better understand the present solution and do not constitute a limitation of the present application; wherein:
[0012] Figure 1 is a schematic diagram of an exemplary system architecture in which the various methods and devices described herein may be implemented according to an embodiment of the present application.
[0013] Figure 2 This is a flowchart of a method for determining prompt information according to an embodiment of the present application.
[0014] Figure 3 is a schematic diagram of a page in a target file according to an embodiment of the present application.
[0015] Figure 4 is a schematic diagram of an intelligent dialogue interface according to an embodiment of the present application.
[0016] Figure 5 is a schematic diagram of obtaining relevance according to an embodiment of the present application.
[0017] Figure 6 4 is a flowchart of a training evaluation model according to an embodiment of the present application.
[0018] Figure 7 This is a flowchart of a method for determining prompt information according to another embodiment of the present application.
[0019] Figure 8 It is a block diagram of a prompt information determination device according to an embodiment of the present application.
[0020] Figure 9 Schematically shows a structural block diagram of an artificial intelligence agent according to an embodiment of the present application; and
[0021] Figure 10 A schematic block diagram of an example electronic device that can be used to implement embodiments of the present application is shown. DETAILED DESCRIPTION
[0022] The following description of exemplary embodiments of the present application is made in conjunction with the accompanying drawings, including various details of the embodiments of the present application to facilitate understanding. These details should be considered as merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications may be made to the embodiments described herein without departing from the scope and spirit of the present application. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.
[0023] In the technical solution of this application, the user information involved (including but not limited to user personal information, user image information, user device information, such as location information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) are all information and data authorized by the user or fully authorized by all parties, and the collection, storage, use, processing, transmission, provision, application and application of the relevant data comply with relevant laws, regulations and standards, take necessary confidentiality measures, do not violate public order and good morals, and provide corresponding operation entrances for users to choose to authorize or refuse.
[0024] A large model refers to a deep learning model with large-scale model parameters. Large models typically contain hundreds of millions, billions, tens of billions, hundreds of billions, trillions, or even more than ten trillion model parameters. Large models may include large language models (LLMs), GPTs (Generative Pre-trained Transformers), large visual models, multimodal large models, and the like. The large models involved in the embodiments of this application may be general large models, or they may be expert large models obtained after fine-tuning based on requirements, and the embodiments of this application are not limited to this. An intelligent agent includes a system or entity that can autonomously perceive the environment, make decisions, and perform actions to complete specific tasks. The large model can provide decision support for the intelligent agent and provide the intelligent agent with the ability to perform reasoning analysis and task planning. The intelligent agent uses the analysis results of the large model to execute or optimize its decision-making process. The intelligent agent can integrate multiple large models to handle different types of tasks.
[0025] In application scenarios involving large models and intelligent agents, users or agents can input instructions and associated documents (such as contracts, technical manuals, and research reports) and receive semantic understanding and feedback (such as question-and-answer, creative writing, and summarization) from the large model. The process involves the agent internally concatenating the historical conversation context, user instructions, and document content into a complete input, which the large model then interprets and generates a response. If the total input length does not exceed the large model's maximum acceptable token limit (e.g., 128k tokens), the entire input can be directly fed into the model, ensuring lossless information transmission. However, the current maximum input length of some large models is typically under 100,000 tokens, while real-world documents (such as technical manuals, legal documents, and novels) can often be millions of tokens long. Therefore, it is necessary to address the difficulty of efficiently extracting document content related to instructions when the input exceeds the large model's maximum input limit, while preserving the document content, structured data, and multimodal information associated with the instructions as completely as possible.
[0026] Within the maximum input limits of large models, extracting instruction-related file content, structured data, and multimodal information faces multiple difficulties. For example, instruction types vary greatly. In addition to needle-in-a-haystack question-answering tasks (such as asking questions about a specific paragraph in a massive text), full-text tasks (such as full-text creation, file comparison, and review) face even greater difficulties. For example, the specific content associated with instructions varies significantly between different intents; identifying content that closely matches instructions is difficult, and the key core parts that need to be retained may be scattered across multiple parts or multiple files within a document. Some answers rely on the large model's certain reasoning capabilities, combined with its understanding of instructions and judgment of file content, to locate them. Some tasks even require reading the entire text, which may exceed the maximum input limits of the large model.
[0027] For example, indiscriminate compression of content in files that is not related to instructions by directly truncating or randomly discarding content, extractive summarization, generative summarization, fixed blocking, etc. may lead to the loss of key information (such as disclaimers in contract terms, key parameters of experimental data, etc.), omission of query-related details (such as specific data, terminology, etc.), and problems such as context fragmentation and repeated calculations; through keyword matching methods, the file content is filtered by rule-based instruction keyword filtering, inverted index retrieval methods, etc., which may lack semantic understanding (such as synonyms, context associations, etc.) and cannot dynamically adapt to the maximum input limit of the large model; through retrieval-augmented generation (RAG), file fragments related to instructions are retrieved and combined with the large model to output answers. However, it relies on a pre-built knowledge base to adapt to the knowledge question and answer scenario, and is not suitable for processing long content in one or more files uploaded by users. Among them, Retrieval Augmented Generation (RAG) refers to enhancing the input of the large model by retrieving more reference information from external knowledge bases or indexes based on the large model, and ultimately generating more accurate and richer answers or content.
[0028] In addition, RAG relies on local fragments and lacks global semantic understanding. For example, it usually retrieves the Top-K document fragments (such as paragraphs or sentences) based on vector similarity, but it cannot guarantee coverage of the global information of the file. For example, in a full-text question answering task, if the answer is scattered across multiple chapters (such as "Contract Termination Conditions" included in Clauses 3.1, 5.2 and Appendix B), RAG may only return some fragments, resulting in incomplete answers. There is a lack of coherence between the fragments retrieved by RAG (such as skipping the middle paragraphs in the logical chain), resulting in inconsistent generated content. For example, when analyzing technical documents, if only "fault symptoms" are retrieved and "root cause analysis" is ignored, the model may incorrectly attribute the cause of the problem. In addition, RAG has certain difficulties in processing structured and non-text content, and it is difficult to process structural layout and multimodal information because it is usually optimized for pure text and it is difficult to effectively process multimodal content such as tables, charts, formulas, audio and video, as well as structural layouts with mixed arrangements between multimodal content. For example, in financial report analysis, key data in a table (such as quarterly revenue growth rate) may be ignored because it is not textualized. The user asks " Figure 2When asked, "What does the trend of the curve in the image indicate?", RAG may generate speculative answers based solely on text descriptions. For example, when a technical drawing is accompanied by text descriptions, RAG's text retrieval mechanism cannot associate the image semantics, resulting in the generated answer being off-base. RAG relies on splicing retrieval fragments, making it difficult to generate content that covers the core ideas of the entire text. For example, in a full-text summarization task, if the document topics are scattered (such as a scientific research paper containing multiple innovative points), the summary generated by RAG may only reflect local highlights and omit the overall contribution. To a certain extent, there is a contradiction between RAG's efficiency and accuracy. For example, to improve retrieval coverage, the retrieval scope needs to be expanded (such as increasing the Top-K value), but this will significantly increase computational costs and may introduce noise (such as low-relevance fragments). RAG usually relies on pre-built file indexes, and the construction time is highly correlated with the length of the text. Index construction time for very long files may be greatly increased.
[0029] Based on at least one of the above problems, an embodiment of the present application proposes a method for determining prompt information, which obtains multiple element blocks by retaining the content blocks and structural layout information in the target file, and accurately determines multiple correlations between the task description information and the multiple element blocks based on the content blocks and their associated structural layout information. It can identify file content that is strongly related to the task intent from a global perspective, realize context association during the compression process, and maintain semantic coherence and task adaptability at the same time, thereby improving the integrity of key content, structural layout information and other data in the prompt information.
[0030] The technical solutions provided by this application will be described in detail below with reference to the accompanying drawings and specific embodiments.
[0031] Figure 1 is a schematic diagram of an exemplary system architecture in which the various methods and devices described herein may be implemented according to an embodiment of the present application.
[0032] It should be noted that Figure 1 What is shown is merely an example of a system architecture to which the embodiments of the present application can be applied, to help those skilled in the art understand the technical content of the present application, but does not mean that the embodiments of the present application cannot be used in other devices, systems, environments or scenarios.
[0033] like Figure 1 As shown, the system architecture 100 according to this embodiment may include a terminal device 102, a network 103, and a server 104. The network 103 is used to provide a medium for a communication link between the terminal device 102 and the server 104. The network 103 may include various connection types, such as wired and / or wireless communication links, etc.
[0034] A user can use a terminal device 102 to interact with a server 104 via a network 103. The user can send task description information 101 and a target file 106 through the interactive interface provided by the terminal device 102. The terminal device 102 can then send the task description information 101 and the target file 106 to the server 104 via the network 103, causing the server 104 to process the task description information 101 and the target file 106 and output prompt information 105. The server 104 then sends the prompt information 105 to the terminal device 102, causing the terminal device 102 to display the prompt information 105 to the user. Alternatively, the server 104 can invoke a large model (or agent) to input the prompt information 105, obtain the large model's response, and send it to the terminal device 102, causing the terminal device 102 to display the response to the user.
[0035] Various communication client applications may be installed on the terminal device 102, such as intelligent assistant applications, knowledge reading applications, web browser applications, search applications, instant messaging tools, email clients, and / or social platform software (for example only). A user may enter task description information 101 and target file 106 in the interactive interface of these client applications, and these client applications will display generated prompt information 105 to the user.
[0036] The terminal device 102 can be configured as various electronic devices with a display screen and supporting web browsing, including but not limited to smart phones, tablet computers, laptop computers, desktop computers, etc.
[0037] Server 104 can be a server that provides various services, such as a backend management server (for example only) that supports content viewed by users through the interactive interface of terminal device 102. The backend management server can invoke an agent to perform data queries in response to received query requests, and feed the query results back to terminal device 102, displaying them through the interactive interface. For example, server 104 can be a standalone physical server, a server cluster or distributed system consisting of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud computing, network services, and middleware services.
[0038] It should be noted that the prompt information determination method provided in the embodiment of the present application can generally be executed by the server 104. Accordingly, the prompt information determination device provided in the embodiment of the present application can also be set in the server 104. The prompt information determination method provided in the embodiment of the present application can also be executed by a server or server cluster that is different from the server 104 and can communicate with the terminal device 102 and / or the server 104. Accordingly, the prompt information determination device provided in the embodiment of the present application can also be set in a server or server cluster that is different from the server 104 and can communicate with the terminal device 102 and / or the server 104.
[0039] Alternatively, the terminal device 102 can install a small model with a small number of model parameters, such as around 7 billion to 14 billion, which is only an example. The small model refers to a large model with relatively small model parameters, such as a large-scale language model (LLM), GPT (Generative Pre-trained Transformer), a large visual model, a large multimodal model, etc. The small model is deployed locally on the terminal device 102, so that the user can send the task description information 101 and the target file 106 through the interactive interface provided by the terminal device 102. The terminal device 102 can execute the prompt information determination method provided in the embodiment of the present application, process the task description information 101 and the target file 106 and output the prompt information 105, or the terminal device 102 can call the small model (or a local agent built based on the small model) to input the prompt information 105, obtain the reply content of the large model and display it to the user. Accordingly, the prompt information determination device provided in the embodiment of the present application can generally be set in the terminal device 102.
[0040] It should be understood that Figure 1 The number of terminal devices and servers in the embodiment is merely illustrative. Any number of terminal devices and servers may be used as required.
[0041] Figure 2 This is a flowchart of a method for determining prompt information according to an embodiment of the present application.
[0042] like Figure 2 As shown, the method 200 may include:
[0043] In operation S210 , multiple correlations between the task description information of the target object and multiple element blocks of the target file are determined. The element blocks include content blocks separated from the target file according to the file structure layout and structural layout information associated with the content blocks.
[0044] Task description information can include information in one or more modalities, such as text, voice, or video, that describes the task content. For example, the target object can include a user, an agent, a terminal device, or a server. The task description information can be determined based on information input by the user, or by receiving information sent by the terminal device or server based on manual operation, event triggering, or scheduled tasks. For example, the task description information can be information about a question raised by the user. The question information can be a query, question, or requirement raised by the user. The question information can be in the form of text or other modalities such as voice, image, etc.
[0045] Relevance is a metric that measures the degree of correlation between task description information and element blocks. For example, multiple element blocks can represent the relationships and structural layout of file content, thereby eliminating contextual disconnection to a certain extent based on the content blocks and their structural layout information. Natural language processing technology can be used to determine relevance to task description information from multiple perspectives and scopes based on the central idea, overall semantics, sentiment, and intent of the target file's content.
[0046] The target file includes a file whose content can be parsed, and may include formats such as files, pictures, and web pages. The target file is, for example, a file uploaded by a user, or a file searched through a specified data source, and there may be one or more. The content block includes the content part with a certain degree of independence in the target file and the associated information such as the position of the content part, which can be obtained by splitting from the target file in combination with the structural layout information. The file structure layout includes the arrangement structure of the content in the target file, which can be determined, for example, according to paragraphs, semantics, typesetting style, and outline structure, such as including two main headings, each of which includes three subheadings, and each subheading includes corresponding content. For example, a paragraph under one of the subheadings can be used as a content block, and its associated structural layout information includes the subheading and the subheading's parent main heading.
[0047] In operation S220 , at least one target element block is screened out from the plurality of element blocks based on the plurality of correlations.
[0048] The target element block includes element blocks that are highly relevant to the task description information, which are selected from multiple element blocks. For example, in a legal contract review scenario, if a user asks "What are the conditions for Party A to unilaterally terminate the cooperation?", multiple target element blocks containing "termination clauses" can be identified from different chapters of the contract to avoid retaining only the first paragraph of the clause description but omitting subsequent exceptions (such as force majeure clauses). Or, in a multimodal scenario, if a user asks " Figure 3What failure mode does the peak in the image correspond to? If the target file contains key graphics (such as technical drawings and data curves), one or more target element blocks containing text and technical drawings that are highly relevant to the task description can be identified. Through cross-modal association capabilities, this avoids compressing only text and ignoring images, which makes it difficult for subsequent large models to understand the semantics of the combined text and graphics.
[0049] In operation S230 , prompt information for the large model is determined based on the task description information and at least one target element block, where the prompt information is used to guide the large model to generate response content for the task description information.
[0050] Prompt engineering is a systematic approach to designing, optimizing, and managing interactive prompts for large models. It aims to guide models to produce high-quality, consistent outputs through precise language construction. Prompts are generated based on the output of prompt engineering and are used to guide large models in generating specific responses.
[0051] For example, a user uploads target files such as marketing reports and sales record files to the agent application and proposes "summarizing the marketing strategy of product A in region B over the past year" as the task description information. The agent can call a tool to first split these files into multiple element blocks according to the file structure layout. Each element block contains a content block (such as a description of a marketing campaign) and structural layout information (such as the page number, chapter, title, etc. of the file where the content is located). Next, the agent can call another tool to determine the relevance between the task description information and each element block. For example, the chapter on region B in a marketing report describes the marketing strategy of product A. Combined with the structural layout information containing the information of this chapter, a content block under this chapter may be describing the marketing strategy of product A over the past year, so the corresponding element block has a high relevance. Then, the element blocks with high relevance are selected as target element blocks, such as several related chapters in the above-mentioned marketing report. The task description information and the target element block content are integrated to construct prompt information, such as "Based on the chapter "XXXXX" about region B in the marketing report, explain the marketing strategy of product A in region B over the past year". The intelligent agent sends this prompt information to the big model, so that the big model generates corresponding reply content based on this prompt information and feeds it back to the user.
[0052] According to an embodiment of the present application, multiple element blocks are obtained by retaining the content blocks and structural layout information in the target file, and multiple correlations between the task description information and the multiple element blocks are accurately determined based on the content blocks and their associated structural layout information. This allows for identifying file content that is strongly related to task intent from a global perspective, achieving contextual association during the compression process while maintaining semantic coherence and task adaptability, thereby improving the integrity of data such as key content, structural layout information, etc. in the prompt information.
[0053] Figure 3 is a schematic diagram of a page in a target file according to an embodiment of the present application.
[0054] In some embodiments, the target file can be split according to the file structure layout to determine multiple content blocks; based on the hierarchical relationship between the outline titles represented by the file structure layout, multiple structural layout information associated with the multiple content blocks can be extracted; based on the file content, file content location, at least one of the target file information of the multiple content blocks, and multiple structural layout information, multiple element blocks can be obtained.
[0055] For example, multiple content blocks can be divided based on chapter divisions, paragraph levels, and content attribution. The hierarchical relationship between outline headings represents the hierarchical association between headings used to present the hierarchical structure of the document's content. For example, a first-level heading contains a second-level heading, and a second-level heading may contain a third-level heading. This superior-subordinate relationship is a hierarchical relationship. If a content block in a document falls under a second-level heading, its structural layout information includes the hierarchical position of the second-level heading in the document and the parent heading of the second-level heading.
[0056] Each element block can be derived based on at least one of the following: the file content, file content location, and target file information contained in the corresponding content block, as well as structural layout information. The file content location refers to the specific location of the content block's file content within the target file, such as the page number or paragraph number within the file. The target file information includes attribute information about the target file itself, such as the file name, file type, and creation time.
[0057] Reference Figure 3 For example, if a page in the target document is displayed, it can be identified according to the rules of title level and paragraph separation, such as obtaining a first-level title 301, a first and second-level title 302, a first paragraph 303, a second paragraph 304, a picture 305, a picture number 306, a second and second-level title 307, and a third paragraph 308. Further, based on these identification results, one or more element blocks are determined.
[0058] For example, each paragraph may correspond to multiple titles at different levels, and each title may have multiple paragraphs. When forming an element block, each paragraph (i.e., file content) is spliced with all the titles (i.e., structural layout information) containing it, and each title (i.e., file content) is spliced with all the paragraphs (i.e., structural layout information) contained in its current level. For example, the first paragraph 303 corresponds to the first-level title 301 and the first and second-level titles 302, and its corresponding element block is represented by para: {"title": "First-level title 301, first and second-level titles 302", "content": "Content of first paragraph 303", "para_index": "Serial number of first paragraph 303", "filename: "Target file name", "document_index": "Serial number of target file"}; the first and second-level titles 302 contain the first paragraph 303 and the second paragraph 304, and its corresponding element block is represented by title: {"title_content": "First and second-level titles 302", "para_content": "First paragraph 303, second paragraph 304", "title_index": "Serial number of first and second-level titles 302", "filename: "Target file name", "document_index": "Serial number of target file"}. If the user uploads multiple target files, the filename and document_index of each content block are filled in according to the file to which they are located.
[0059] According to an embodiment of the present application, by splitting the target file into content blocks and obtaining relevant structural layout information to form element blocks, the contextual relevance can be enhanced, the information dimension can be enriched, the relationship between the various parts of information within the file can be better reflected, and it is helpful to characterize the relationship between the file content and the task description information.
[0060] In some embodiments, the target file is split according to the file structure layout, and the determination of multiple content blocks includes: obtaining multimodal context information based on at least one of the file structure layout, semantic associations between multimodal file contents, keyword associations between multimodal file contents, and logical associations between multimodal file contents; and splitting the target file based on the multimodal context information to obtain content blocks including multimodal file contents.
[0061] Multimodal file content refers to the different types of data contained in the target file, such as text, tables, formulas, pictures, audio and video. Semantic association includes the connection between the meanings of different modal file contents, such as Figure 3 There is a text description " Figure 2 Describes the beautiful scenery of the seaside..." and the corresponding Figure 2Keyword association refers to the connection formed by the same or related keywords in the content of different modal files, such as Figure 3 Both the Chinese text description and the picture title have " Figure 2 ". Logical associations include the relationship between the contents of different modal files at the logical level, such as cause and effect, sequence, etc., such as the logical association between the second paragraph 304, the picture 305, and the picture sequence number 306 in terms of semantics (such as text semantics and image semantics identified by the picture) and the logical association in terms of layout. Multimodal context information includes information formed by comprehensively considering the semantic, keyword, and logical associations between the contents of multimodal files. For example, Figure 3 The file content of the content block in the page shown includes a second paragraph 304 , a picture 305 , and a picture sequence number 306 .
[0062] According to the embodiments of the present application, by considering the various associations between multimodal document content, content blocks can contain information closely related to a specific topic in multiple modalities, achieving multimodal contextual association, improving the integrity of key content, structural layout information, structured content, and other data in the prompt information, effectively guiding the semantics of the large model to associate multimodal content, and preventing the generated response content from deviating from reality. The structured content includes the structured information of the mutual associations between multimodal content.
[0063] Figure 4 is a schematic diagram of an intelligent dialogue interface according to an embodiment of the present application.
[0064] In some embodiments, first description information input by the target object can also be obtained; based on a preset rewriting strategy, the first description information is rewritten to obtain second description information, and the rewriting strategy includes integrating at least one of the dialogue context of the target object, target language translation, and pronoun rewriting; task description information is obtained based on the first description information and the second description information.
[0065] The first description information can be a command entered by the user during a conversation. "Integrating the target subject's conversation context" is used to rewrite the first description information based on the target subject's previous conversational content. "Target language translation" is used to translate the first description information into another language. "Pronoun rewriting" is used to convert the pronouns in the first description information. For example, a rewriting model (either lightweight or large) can be used to process the first description information and the conversation context based on the rewriting strategy to generate the second description information. The first and second description information are then concatenated to obtain the task description information.
[0066] like Figure 4As shown, the agent application deployed on terminal device 300 can display an intelligent dialogue interface. This interface displays an interactive area 310, including an "Add Attachment" function 311, a "Send Image" function 312, and a "Send" button 313. Users can enter commands in the input box of interactive area 310, for example, entering first descriptive information 320, such as "Write a product promotion copy," and uploading a product introduction file 330 through the "Add Attachment" function 311. Users can also view historical conversations in the historical conversation area 340 (which has been stored in the backend).
[0067] For example, a user enters the first description "Write a promotional copy for our product." The agent processes this information based on a pre-set rewriting strategy. For example, if the previous conversation mentioned product A as featuring natural ingredients, suitable for people with sensitive skin, and targeted at customers in Region C, the agent then integrates the conversation context, target language translation, and pronoun rewriting strategies to generate a second description in Region C's language: "Write a promotional copy for our product A, featuring natural ingredients and suitable for people with sensitive skin." "Product A" is a rewritten pronoun for "product" in the first description.
[0068] According to an embodiment of the present application, through a rewriting strategy, the possible task intention of the target object can be generalized in combination with the first description information and the second description information, so as to match a target element module that meets the intention.
[0069] In some embodiments, determining multiple correlations between the task description information of the target object and multiple element blocks of the target file includes: determining multiple element features of the multiple element blocks based on at least one of the task description information, the content block and the structural layout information; inputting the multiple element features into the evaluation model to obtain multiple correlations output by the evaluation model, wherein the parameter amount of the evaluation model is smaller than the parameter amount of the large model.
[0070] For example, for text content blocks, multiple element features can be obtained by extracting keywords as element features, extracting the file level and page number of the element block from the structural layout information, and extracting element features through semantic analysis by combining the task description information and the content block. The evaluation model is used to evaluate the relevance between the task description information and the element block based on the input element features. Compared to large models, the evaluation model is lightweight. The parameter count refers to the number of learnable parameters in the model, which reflects the model's complexity and size.
[0071] For example, the evaluation model can be a shallow neural network model, a vector machine model, or a large language model with smaller parameters than a large model. For example, RankSVM (Ranking Support Vector Machine) uses the Learning to Rank (LTR) algorithm to solve the relevance ranking problem of multiple element blocks.
[0072] Exemplarily, a parallel strategy can be adopted to use the evaluation model to simultaneously calculate the correlation between the task description information and at least two element blocks, thereby maintaining performance while improving the effect, which can better play a role in intelligent agent applications and improve end-to-end response speed.
[0073] According to the embodiments of this application, element features are extracted from multiple aspects of task description information, content blocks, and structural layout information, enabling the evaluation model to more accurately predict relevance. This lightweight evaluation model reduces computing resource consumption and processing time while maintaining accurate relevance assessment, enabling rapid evaluation of the relevance of a large number of element blocks with task description information.
[0074] In some embodiments, determining multiple element features of multiple element blocks based on at least one of task description information, content blocks and structural layout information includes: obtaining multiple element features based on multiple similarities between the task description information and the multiple element blocks, the ranking of multiple similarities, multiple keyword overlaps between the task description information and the multiple element blocks, the ranking of multiple keyword overlaps, the file content positions of the multiple element blocks, the structural layout information of the multiple element blocks, and at least one of the semantic information of the content blocks of the multiple element blocks.
[0075] For example, an element block is represented as para:{"title":"one or more layers of outline titles","content":"paragraph content","para_index":"paragraph number","filename":"target file name","document_index":"target file number"}. Element features include at least one of similarity features, sorting features, vocabulary features, semantic features, and outline features.
[0076] Figure 5 FIG. 1 is a schematic diagram of obtaining relevance according to an embodiment of the present application. Figure 5, similarity can be obtained by calculating the Euclidean distance between the vector of the task description information and the vector of the element block (only as an example), thereby obtaining similarity features, the vector of the element block is represented by the vector of at least one of the content of title, content, para_index, filename and document _index; the ranking features are obtained by sorting the similarity corresponding to the element block in multiple similarities, sorting the corresponding keyword overlap in multiple keyword overlaps and sorting the file content position in the corresponding content block; for example, by segmenting the task description information and extracting the vocabulary, using all the segmented words to build a dictionary (where low-frequency words can be filtered), and then calculating the number of overlaps between the words in the dictionary and each word in each element block, the keyword overlap is obtained, and then the vocabulary features are obtained; semantic information is extracted from the content in the content block to obtain semantic features; outline features are obtained based on the title and hierarchical relationship in the title. It can be understood that element features are not limited to Figure 5 The content shown, for example, can also be used to obtain position features and target file features based on para_index, filename, and document_index. Then, the element features of each element block are input into the evaluation model 510 to obtain multiple relevance levels.
[0077] According to the embodiments of the present application, element block features are obtained through similarity, various types of sorting, file content position, keyword overlap and semantic information, which can retain data such as overall semantics, intention, structural information, etc., making it easier to determine the relevance with task description information from multiple angles and multiple ranges.
[0078] In some embodiments, based on multiple relevance, filtering out at least one target element block from multiple element blocks includes: based on multiple relevance, filtering out at least one target element block that meets a preset length from multiple element blocks, the preset length being pre-set according to the maximum number of input tokens of the large model; wherein the large model is predetermined from multiple candidate large models, and the maximum number of input tokens of at least two of the multiple candidate large models is different.
[0079] For example, the preset length can be less than or equal to the maximum number of input tokens of the large model. Multiple candidate large models may include large language models (LLM), GPT (Generative Pre-trained Transformer), visual large models, multimodal large models, and other types of models, or include models of the same type but with one or more different architectures, number of parameters, training data, etc. Different candidate large models may have different maximum numbers of input tokens, that is, the maximum number of tokens that can be processed. For example, based on the content order in the target file, the target element blocks are selected in sequence until the preset length is reached.
[0080] In some embodiments, based on factors such as computing resources or task scenarios, the preset length can be determined by further narrowing the range within the maximum number of input tokens of the large model. The preset length can be dynamically adjusted based on factors such as computing resources or task scenarios.
[0081] According to embodiments of the present application, the maximum number of input tokens for different models or the same model can be dynamically adapted based on relevance, selecting at least one target element block within a preset length limit (e.g., compressible to 128k or further compressed to 60k tokens). This allows for dynamic retention of more relevant content for different compression targets (i.e., preset lengths), extracting file content that meets the model length limit while maintaining semantic coherence and task suitability.
[0082] In some embodiments, based on the task description information and at least one target element block, determining the prompt information for the large model includes: when multiple target element blocks are screened out, based on the file content positions of the respective content blocks of the multiple target element blocks, sequentially splicing multiple file contents to obtain a splicing result; based on the task description information and the splicing result, obtaining the prompt information of the large model.
[0083] Sequential splicing involves splicing the target element blocks sequentially, according to their file content location. The splicing result includes the entire content resulting from sequentially splicing the target element blocks. The task description and the splicing result can then be combined to form a prompt.
[0084] According to an embodiment of the present application, by splicing the target element block content in the order of the file content position, the information provided to the large model has inherent logical coherence, which facilitates the large model to reason and generate based on more ordered information, thereby improving the logic and coherence of the generated content.
[0085] In some embodiments, based on the task description information and at least one target element block, determining the prompt information for the large model includes: when the content block of at least one target element block includes multimodal file content, adding multimodal output instructions to the prompt information of the large model to guide the large model to generate multimodal response content of the task description information.
[0086] For example, multimodal content includes at least one of the following: an accessible link to an image obtained by uploading the image in the target file to a cloud storage service; table content and table identifiers obtained by parsing a table in the target file; a formula represented in a specified description language (such as LaTeX) obtained by parsing the formula in the target file; an audio text description obtained by performing speech recognition on the audio in the target file; and a video text description obtained by performing at least one of visual analysis, subtitle recognition, and speech recognition on the video in the target file.
[0087] The multimodal output instructions guide the large model to generate content that is not present in the target file or can be the original text of the target file. If the task description information determines that the original text needs to be referenced, the multimodal output instructions can also include at least one of an accessible link indicating an image, a table identifier, a formula representing a descriptive language, an audio text description, and a video text description, so that the content output by the large model can correctly reference the original text.
[0088] According to the embodiments of the present application, multimodal compatibility of the target file content is achieved in the prompt information, the multimodal structured content is jointly retained, and multimodal output instructions can be added accordingly. This allows the large model to consider multimodal information and output the response content in a multimodal form that is more in line with the target object's intention.
[0089] The following further describes the process of training and evaluating the model.
[0090] Figure 6 4 is a flowchart of a training evaluation model according to an embodiment of the present application.
[0091] like Figure 6 As shown, the training evaluation model includes:
[0092] In operation S610 , a plurality of correlation labels are obtained between the task description sample and a plurality of element block samples of the file sample. The element block samples include content block samples separated from the file sample according to the file structure layout and structural layout information associated with the content block samples.
[0093] Exemplarily, the task description samples, file samples, element block samples, content block samples, predicted relevance, relevance labels, etc. in the training phase are the same or similar to the interpretation and processing methods of the above-mentioned task description information, target files, element blocks, content blocks, relevance, etc. in the reasoning phase, only to reflect the differences in the embodiments.
[0094] In some embodiments, the file sample is, for example, user log data obtained with the user's authorization, or can be combined with a large model to generate and expand the content of the file sample, such as paragraph content, outline title, and multimodal content.
[0095] For example, element block samples can include relevance data samples and prior data samples. Relevance data samples can be semantically related to the task description sample, such as those primarily concerned with questions and answers or requiring the extraction of specific content. Relevance labels are obtained by identifying and splitting the document sample and then processing it based on the large model (e.g., attention weights). Prior data samples can be semantically unrelated, but the large model can use their content to generate responses that better align with the intended purpose. For example, for summarizing, continuing, and generating outlines, prior data samples and relevance labels for different types of task description samples can be generated through manually formulated rules or with the help of large model responses. For example, the prior data sample for summarizing is "summary paragraph + conclusion paragraph + outline," and the prior data sample for continuing is the final paragraph of the document. Relevance labels with high scores are positive examples, while those with low scores are negative examples.
[0096] In operation S620 , a plurality of element features of a plurality of element block samples are obtained based on at least one of the task description sample, the content block sample, and structural layout information associated with the content block sample.
[0097] For example, multiple element features are obtained based on at least one of multiple similarities between the task description sample and the multiple element block samples, the ranking of the multiple similarities, multiple keyword overlaps between the task description sample and the multiple element block samples, the ranking of the multiple keyword overlaps, the file content positions of the multiple element block samples, the structural layout information of the multiple element blocks, and the semantic information of the content block samples of the multiple element blocks.
[0098] In operation S630 , a plurality of element features of the plurality of element block samples are input into an evaluation model, and a plurality of predicted correlations between the task description sample output by the evaluation model and the plurality of element block samples are obtained.
[0099] In operation S640 , an evaluation model is trained based on the difference between the rankings of the plurality of relevance labels and the rankings of the plurality of predicted relevances.
[0100] For example, a ranking loss function, such as Triplet Margin Loss, is used to measure the difference between the relevance label ranking and the predicted relevance ranking. Through the backpropagation algorithm, the gradient generated by this difference is transferred to the various parameters of the evaluation model, and the model parameters are adjusted to gradually reduce the ranking difference.
[0101] According to the embodiments of this application, element features are extracted from multiple aspects, including task description samples, content block samples, and structural layout information, enabling the evaluation model to learn information from multiple dimensions during the training phase. Training based on ranking differences focuses on optimizing the evaluation model's performance in comparing and ranking the relevance of multiple element blocks to task descriptions, enabling more accurate prediction of relevance.
[0102] In some embodiments, multiple relevance labels are obtained by the following operations: using a large model to process a task description sample and multiple element block samples to obtain multiple attention matrices, where the attention matrix is obtained based on at least one of the attention matrix between the corresponding element block sample and the file sample and the attention matrix between the element block sample and the task description sample; and obtaining multiple relevance labels based on at least part of the attention weight of each of the multiple attention matrices.
[0103] For example, the large model includes multiple layers of attention layers stacked together. The relevance label of the element block sample can be determined based on the attention layer score of the large model. It can be obtained by the output of a certain attention layer or by the weighted score of the output of multiple attention layers. The attention layer score can be calculated using the average value of the top 10% elements in the attention weight matrix between the element block sample and the file sample, and the attention weight matrix between the element block sample and the task description sample, and can be directly used as the relevance label. Among them, considering that the weight of special positions (token and itself, first token) is too high, the token weight value of the key area (the area containing tokens in special positions) is weakened, so the key area is normalized and the score is calculated.
[0104] According to the embodiments of the present application, the attention mechanism can capture the importance weight relationship between different parts of the input. By utilizing the attention layer score of the large model to determine the relevance label of the element block sample, it can deeply explore the semantic association between the element block and the task description and the file, and reflect the actual correlation between the element block and the task description information.
[0105] In some embodiments, the effect of compressing responses to different lengths can be tested using a large model. For example, the large model response is used as a pseudo training label, and the complete sequence is [question][document][answer]. Recall is obtained by sorting the response-doc and query-doc attention weight scores, and the effects of compression to different lengths are compared.
[0106] Figure 7 This is a flowchart of a method for determining prompt information according to another embodiment of the present application.
[0107] like Figure 7 As shown, the interface can be encapsulated and allowed to be called externally, such as directly by a user, or called as a tool or service in an agent scenario. Specifically, the method 700 may include:
[0108] In operation 701, a user query (e.g., first description information) and the results of parsing one or more target documents are received. For example, the content is identified and detected by paragraph, and the category, location, and content of each paragraph are obtained. A preset length of the desired prompt information may also be received. This can be preset based on one or more factors, such as the capabilities of the large model and the length of the article.
[0109] In operation S702, pre-processing is performed, such as loading the parameters of the large model to be called, the preset prompt, etc.
[0110] In operation S703, the conversation context is compressed. This compressed conversation context may include one or more historical conversations within the current conversation, as well as information from other conversations or user-specific information. For example, compression may include focusing on the user's most recent (most recent) uploaded file, the most recent conversation, and the current user's query.
[0111] In operation S704, it is determined whether the compression result obtained by executing operation S703 exceeds the preset length. If not, operation S708 is executed using the prompt loaded in the pre-processing. If it has exceeded, operation S705 is executed.
[0112] In operation S705 , the query is rewritten, for example, by invoking a rewriting model to fuse multiple rounds of dialogue, rewriting into Chinese and English, and performing reference rewriting to obtain a rewritten query (such as the second description information).
[0113] In operation S706, the file parsing results are preprocessed. For example, the parsing results are processed into a structure encoding each paragraph and corresponding position. In some embodiments, the parsing results can also be processed into a structure encoding multiple multimodal paragraph contents, paragraph positions, outline level information, target file information, etc. to form element blocks.
[0114] In operation S707, the target element block is selected. First, all file information (file name, with position-coded paragraphs, and multiple element blocks formed by chapter title structure), the original query and the rewritten query, and a preset length are input. Next, feature extraction is performed, such as calculating at least one of similarity features, ranking features, vocabulary features, semantic features, and outline features. The vocabulary features rely on Jieba word segmentation of the original query and the rewritten query, followed by vocabulary matching. Next, a preloaded evaluation model is used for parallel scoring, that is, multiple element blocks are processed in parallel to predict the relevance of each element block. The relevance scores are then sorted. Finally, element blocks are selected one by one from the most relevant to the least relevant until the preset length is exceeded.
[0115] In operation S708, post-processing occurs. For example, the original query, the rewritten query, and all target element blocks are concatenated. The total concatenated token length is checked, and if it exceeds the maximum token length, the concatenated blocks are truncated. All target element blocks are concatenated sequentially based on their file content location. Finally, a prompt message is output, along with the number of characters and tokens in the output text.
[0116] Therefore, by splitting the file, the multimodal relevance of text, images, and other content in the original text is retained during parsing. Then, a trained lightweight model (i.e., evaluation model) is used to score the relevance of the query and all target element blocks before and after the rewrite. The blocks are sorted by score, and the top paragraphs that meet the total length (preset length) requirement are selected. After splicing them together in the order of the original text, the final prompt result is combined with the query before and after the rewrite. When inputting the large model, the corresponding multimodal output instructions are spliced in, so that the model retains multimodal related information such as images and text when outputting the results. This can maintain overall performance while solving possible limitations in other solutions (such as indiscriminate compression, keyword matching, and retrieval enhancement).
[0117] For example, it can support various needs such as related queries, summaries, and creations of ultra-long texts that far exceed the maximum token number input limit of the large model. For example, in some scenarios, it can support reading of file contents of over 2 million words, a maximum of 200M, and a maximum of 100 articles. However, the technical solution of this application is not limited to these indicators. For example, dynamic adaptation can be achieved by adjusting the preset length.
[0118] According to the embodiments of the present application, through the query-guided related content compression method, it is possible to dynamically identify content that is strongly related to the target object's intention, obtain prompt information that meets the preset length limit, and at the same time maintain semantic coherence and task adaptability (such as question-and-answer, summary or creation tasks), thereby better achieving the advantages of multi-scenario compatibility, complete retention of relevant content, adaptation to the target compression length (i.e., preset length), multimodal content compatibility, maintaining semantic coherence and strong scalability.
[0119] Among them, multi-scenario compatibility means that it can adapt to various types of queries (full-text creation, document comparison, review, etc.), and also adapt to different business scenarios, such as direct calls by users, or calls as tools or services in intelligent agent scenarios, because its interface is simple and easy to understand, the input is query and document, and the output is a prompt for requesting the model, which is easy to understand and use for intelligent agents or as a separate interface call; full retention of relevant content means dynamically identifying highly relevant content (such as text paragraphs, tables, chart descriptions, etc.) based on the semantics of task description information, and retaining it as completely as possible instead of uniformly deleting it; target compression length adaptation means that the compressed text strictly complies with the model input length limit, and supports dynamic and flexible adjustment of the preset length of target compression; multimodal content compatibility means that in multimodal mixed files, for example, the synchronous retention of text descriptions and associated images (for example, retaining "see" Figure 2 " directional statements and corresponding picture titles); maintaining semantic coherence means avoiding logical breaks caused by compression (such as retaining the complete reasoning process of the causal chain "hypothesis A → deduction B → conclusion C"); strong scalability means that new file samples can be obtained through continuously updated logs to train the evaluation model, so as to continuously upgrade and expand to support new query demand types.
[0120] Figure 8 It is a block diagram of a prompt information determination device according to an embodiment of the present application.
[0121] like Figure 8 As shown, the prompt information determination device 800 may include a relevance module 810 , a screening module 820 and a prompt information module 830 .
[0122] The relevance module 810 may perform operation S210 to determine multiple relevance between the task description information of the target object and multiple element blocks of the target file, where the element blocks include content blocks split from the target file according to the file structure layout and structural layout information associated with the content blocks.
[0123] The screening module 820 may perform operation S220 for screening out at least one target element block from the multiple element blocks based on the multiple correlations.
[0124] The prompt information module 830 may perform operation S230 to determine prompt information for the large model based on the task description information and at least one target element block, where the prompt information is used to guide the large model to generate response content for the task description information.
[0125] In some embodiments, the relevance module may further include a first determining subunit and an evaluation unit. The first determining subunit is configured to determine multiple element features of the multiple element blocks based on at least one of the task description information, the content block, and the structural layout information; and the evaluation unit is configured to input the multiple element features into an evaluation model and obtain multiple relevances output by the evaluation model, wherein the number of parameters of the evaluation model is smaller than that of the large model.
[0126] In some embodiments, the first determination subunit is also used to obtain multiple element features based on at least one of multiple similarities between the task description information and multiple element blocks, the ranking of multiple similarities, multiple keyword overlaps between the task description information and multiple element blocks, the ranking of multiple keyword overlaps, the file content positions of the multiple element blocks, the structural layout information of the multiple element blocks, and the semantic information of the content blocks of the multiple element blocks.
[0127] In some embodiments, the apparatus 800 further includes a target file splitting module and a structure layout information extraction module, configured to obtain a plurality of element blocks based on at least one of the file content, file content location, and target file information of each of the plurality of content blocks, as well as a plurality of structure layout information. The target file splitting module is configured to split the target file according to the file structure layout to determine a plurality of content blocks; the structure layout information extraction module is configured to extract a plurality of structure layout information associated with the plurality of content blocks based on the hierarchical relationship between outline headings represented by the file structure layout.
[0128] In some embodiments, the target file splitting module includes a first acquisition subunit and a second acquisition subunit. The first acquisition subunit is configured to obtain multimodal context information based on at least one of the following: file structure layout, semantic associations between multimodal file contents, keyword associations between multimodal file contents, and logical associations between multimodal file contents; and the second acquisition subunit is configured to split the target file based on the multimodal context information to obtain content blocks including the multimodal file contents.
[0129] In some embodiments, the prompt information module further includes a splicing unit and a prompt information generating unit. The splicing unit is used to sequentially splice the contents of the multiple files based on the file content positions of the respective content blocks of the multiple target element blocks when multiple target element blocks are screened out, to obtain a splicing result; and the prompt information generating unit is used to obtain prompt information of the large model based on the task description information and the splicing result.
[0130] In some embodiments, the prompt information module is also used to add multimodal output instructions to the prompt information of the large model when the content block of at least one target element block includes multimodal file content, so as to guide the large model to generate multimodal response content of the task description information.
[0131] In some embodiments, the screening module is also used to screen out at least one target element block that meets a preset length from multiple element blocks based on multiple relevances, and the preset length is pre-set according to the maximum number of input tokens of the large model; wherein the large model is predetermined from multiple candidate large models, and the maximum number of input tokens of at least two of the multiple candidate large models is different.
[0132] In some embodiments, the apparatus 800 further includes an input acquisition module, a rewriting module, and a task information acquisition module. The input acquisition module is configured to acquire first description information input by a target object; the rewriting module is configured to rewrite the first description information based on a preset rewriting strategy to obtain second description information, wherein the rewriting strategy includes at least one of integrating the target object's conversation context, target language translation, and pronoun rewriting; and the task information acquisition module is configured to obtain task description information based on the first description information and the second description information.
[0133] In some embodiments, the device 800 may also include a training module for obtaining multiple correlation labels between the task description sample and multiple element block samples of the file sample, the element block samples including content block samples split from the file sample according to the file structure layout and structural layout information associated with the content block samples; obtaining multiple element features of the multiple element block samples based on the task description sample, the content block sample and at least one of the structural layout information associated with the content block samples; inputting the multiple element features of the multiple element block samples into the evaluation model to be trained, and obtaining multiple predicted correlations between the task description sample output by the evaluation model and the multiple element block samples; and training the evaluation model based on the difference between the ranking of the multiple correlation labels and the ranking of the multiple predicted correlations.
[0134] In some embodiments, multiple relevance labels are obtained by the following operations: using a large model to process a task description sample and multiple element block samples to obtain multiple attention matrices, where the attention matrix is obtained based on at least one of the attention matrix between the corresponding element block sample and the file sample and the attention matrix between the element block sample and the task description sample; and obtaining multiple relevance labels based on at least part of the attention weight of each of the multiple attention matrices.
[0135] For the parts not mentioned in the apparatus part, they can be understood with reference to the various embodiments of the above-mentioned method. That is, the apparatus part includes modules for executing the various steps of any one of the method embodiments described above. In addition, the implementation methods, technical problems solved, functions achieved, and technical effects achieved of each module / unit / subunit, etc. in the apparatus part embodiment are respectively the same or similar to the implementation methods, technical problems solved, functions achieved, and technical effects achieved of each corresponding step in the method part embodiment, and will not be repeated here.
[0136] Figure 9 The structural block diagram of an artificial intelligence agent according to an embodiment of the present application is schematically shown.
[0137] In the embodiments of the present application, inspired by the von Neumann structure in modern computer theory, such as Figure 9 As shown, the intelligent agent 900 may include multiple core modules: an input module 910 , a control module 920 and an output module 930 .
[0138] In the example, the input module 910 is responsible for receiving or perceiving information such as calls, queries, requests, instructions, signals, question information or data from the outside world (e.g., a user or an external environment), and converting it into a format that the agent 900 can understand and process. The input module 910 is the primary link for the agent 900 to interact with the outside world. It enables the agent 900 to efficiently and accurately obtain the necessary "sensory" information from the outside world and respond to this information. In the example, the input module 910 can be used to receive input information. The input information may include prompt information as described above.
[0139] In this example, processing module 920 is the core support for agent 900's ability to handle complex tasks. Processing module 920 is used to determine a generation task based on the input information received by input module 910. Based on the generation task, a macro model is determined. By invoking the macro model, the input information is processed to generate output information.
[0140] In an example, the output module 930 may be configured to output the output information obtained by the processing module 920. The output information may include the response content to the task description information as described above.
[0141] In an example, the processing module 920 may include a control unit 921 , a storage unit 922 , and an operation unit 923 .
[0142] During operation, the control unit 921 will continuously interact with the storage unit 922, the computing unit 923, and / or the output module 930. However, in the embodiment of the present application, the control unit 921 acts as a single initiator to initiate communication with the storage unit 930, the computing unit 923, and / or the output module 930, and there may be no communication coupling between the storage unit 922, the computing unit 923, and the output module 930.
[0143] In this example, the performance of the control unit 921 may be closely related to the large model underlying the agent 900. To fully utilize the capabilities of the large language model, the internal structure of the control unit 921 may be designed to be highly configurable and extensible to cope with various types of tasks and requirements in real-world scenarios.
[0144] The storage unit 922 may be responsible for memorizing information such as historical conversations, event flows, etc. The aforementioned prompt information and the reply content to the task description information, etc. may all be included in the storage unit 922 .
[0145] In this example, after receiving input information, agent 900 can use the input information to retrieve target files from storage unit 922, determine relevance, filter target elements, and obtain prompt information, which is then fed back to control unit 921. Control unit 921 can then invoke the large model to process the prompt information and generate a response to the task description information.
[0146] The operation unit 923 can be regarded as a predefined tool library, such as a format conversion tool, which can be included in the operation unit 923 .
[0147] In this example, when AI agent 900 needs to render output data, it can call the relevant renderer and display tools from computing unit 923 and feed them back to processing module 920. Processing module 920 can then use the fed-back renderer and display tools to pass the rendered results to output module 930. It is understandable that although large language models have excellent language understanding and generation capabilities, like humans, they are limited in the tasks they can solve without the help of any tools. Once AI agent 900 is given the ability to call tools, it can perform tasks such as result display.
[0148] The intelligent agent 900 according to the embodiment of the present application can simply and effectively improve the level of intelligence, and enhance flexibility and versatility.
[0149] According to an embodiment of the present application, the present application also provides an electronic device, a readable storage medium and a computer program product.
[0150] According to an embodiment of the present application, an electronic device includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the above method.
[0151] According to an embodiment of the present application, a non-transitory computer-readable storage medium stores computer instructions, wherein the computer instructions are used to enable a computer to execute the above method.
[0152] According to an embodiment of the present application, a computer program product includes a computer program, and the computer program implements the above method when executed by a processor.
[0153] According to an embodiment of the present application, the large model includes a computer program, and the computer program implements the above method when executed by a processor.
[0154] Figure 10 A schematic block diagram of an example electronic device 1000 that can be used to implement embodiments of the present application is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present application described and / or claimed herein.
[0155] like Figure 10 As shown, electronic device 1000 includes a computing unit 1001, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 1002 or a computer program loaded from a storage unit 1108 into a random access memory (RAM) 1003. Various programs and data required for operation can also be stored in RAM 1003. Computing unit 1001, ROM 1002, and RAM 1003 are connected to each other via a bus 1004. An input / output (I / O) interface 1005 is also connected to bus 1004.
[0156] Multiple components in the electronic device 1000 are connected to the I / O interface 1005, including an input unit 1006, such as a keyboard, a mouse, etc.; an output unit 1007, such as various types of displays, speakers, etc.; a storage unit 1008, such as a magnetic disk, an optical disk, etc.; and a communication unit 1009, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 1009 allows the electronic device 1000 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0157] Computing unit 1001 can be any general-purpose and / or specialized processing component with processing and computing capabilities. Some examples of computing unit 1001 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various specialized artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. Computing unit 1001 performs the methods and processes described above. For example, in some embodiments, the methods described above may be implemented as a computer software program tangibly embodied in a machine-readable medium, such as storage unit 1008. In some embodiments, part or all of the computer program may be loaded and / or installed onto device 1000 via ROM 1002 and / or communication unit 1009. When the computer program is loaded into RAM 1003 and executed by computing unit 1001, one or more steps of the methods described above may be performed. Alternatively, in other embodiments, computing unit 1001 may be configured to perform the methods described above by any other suitable means (e.g., via firmware).
[0158] Various embodiments of the systems and techniques described above can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system-on-a-chip systems (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system that includes at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.
[0159] The program code for implementing the methods of the present application can be written in any combination of one or more programming languages. Such program code can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, so that when the program code is executed by the processor or controller, the functions / operations specified in the flow charts and / or block diagrams are implemented. The program code can be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0160] In the context of this application, a machine-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, device, or apparatus. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of machine-readable storage media may include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), optical fibers, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0161] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0162] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer with a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), and the Internet.
[0163] A computer system may include a client and a server. The client and server are generally remote from each other and typically interact through a communication network. The client-server relationship arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other. The server may be a cloud server, a server in a distributed system, or a server integrated with a blockchain.
[0164] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in this application can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this application can be achieved. This is not a limitation herein.
[0165] The above specific embodiments do not constitute a limitation on the scope of protection of this application. Those skilled in the art will appreciate that various modifications, combinations, sub-combinations, and substitutions may be made based on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this application shall be included within the scope of protection of this application.
Claims
1. A method for determining prompt information, comprising: Determining multiple correlations between the task description information of the target object and multiple element blocks of the target file, wherein the element blocks include content blocks separated from the target file according to the file structure layout and structural layout information associated with the content blocks; Based on the multiple correlations, screening out at least one target element block from the multiple element blocks; Based on the task description information and the at least one target element block, prompt information for the large model is determined, where the prompt information is used to guide the large model to generate response content for the task description information.
2. The method according to claim 1, wherein Determining multiple correlations between the task description information of the target object and multiple element blocks of the target file includes: determining a plurality of element features of the plurality of element blocks based on at least one of the task description information, the content block, and the structural layout information; The plurality of element features are input into an evaluation model to obtain the plurality of correlations output by the evaluation model, wherein the number of parameters of the evaluation model is smaller than the number of parameters of the large model.
3. The method according to claim 2, wherein: The determining of the plurality of element features of the plurality of element blocks based on at least one of the task description information, the content block, and the structural layout information includes: The multiple element features are obtained based on at least one of multiple similarities between the task description information and the multiple element blocks, the ranking of the multiple similarities, multiple keyword overlaps between the task description information and the multiple element blocks, the ranking of the multiple keyword overlaps, the file content positions of the multiple element blocks, the structural layout information of the multiple element blocks, and the semantic information of the content blocks of the multiple element blocks.
4. The method according to claim 1, further comprising: Splitting the target file according to the file structure layout to determine a plurality of content blocks; Extracting a plurality of structural layout information associated with a plurality of content blocks according to a hierarchical relationship between outline titles represented by the file structure layout; The plurality of element blocks are acquired based on at least one of the file content, the file content location, and the target file information of each of the plurality of content blocks, and the plurality of structural layout information.
5. The method according to claim 4, wherein The splitting of the target file according to the file structure layout to determine the plurality of content blocks comprises: Obtaining multimodal context information based on at least one of the file structure layout, semantic associations between multimodal file contents, keyword associations between multimodal file contents, and logical associations between multimodal file contents; The target file is split based on the multimodal context information to obtain content blocks including multimodal file content.
6. The method according to claim 1, wherein The determining of prompt information for the large model based on the task description information and the at least one target element block includes: When a plurality of target element blocks are screened out, sequentially splicing the contents of the plurality of files based on the file content positions of the respective content blocks of the plurality of target element blocks to obtain a splicing result; Based on the task description information and the splicing result, prompt information of the large model is obtained.
7. The method according to any one of claims 1 to 6, wherein The determining of prompt information for the large model based on the task description information and the at least one target element block includes: In the case where the content block of the at least one target element block includes multimodal file content, a multimodal output instruction is added to the prompt information of the large model to guide the large model to generate multimodal response content of the task description information.
8. The method according to claim 1, wherein The selecting at least one target element block from the plurality of element blocks based on the plurality of correlations comprises: Based on the multiple relevances, screening out the at least one target element block that meets a preset length from the multiple element blocks, where the preset length is pre-set according to a maximum number of input tokens of the large model; The large model is predetermined from a plurality of candidate large models, and at least two of the plurality of candidate large models have different maximum numbers of input tokens.
9. The method according to claim 1, further comprising: Obtaining first description information input by the target object; rewriting the first description information based on a preset rewriting strategy to obtain second description information, wherein the rewriting strategy includes at least one of integrating the conversation context of the target object, target language translation, and pronoun rewriting; The task description information is obtained based on the first description information and the second description information.
10. The method according to claim 2, wherein: The evaluation model is trained according to the following operations: Acquire multiple correlation labels between the task description sample and multiple element block samples of the file sample, wherein the element block samples include content block samples split from the file sample according to the file structure layout and structural layout information associated with the content block samples; acquiring a plurality of element features of the plurality of element block samples based on at least one of the task description sample, the content block sample, and structural layout information associated with the content block sample; Inputting multiple element features of the multiple element block samples into an evaluation model to be trained, and obtaining multiple predicted correlations between the task description sample output by the evaluation model and the multiple element block samples; The evaluation model is trained based on a difference between the ranking of the plurality of relevance labels and the ranking of the plurality of predicted relevances.
11. The method according to claim 10, wherein: The multiple relevance labels are obtained by the following operations: Processing the task description sample and the plurality of element block samples using the large model to obtain a plurality of attention matrices, wherein the attention matrix is obtained according to at least one of an attention matrix between a corresponding element block sample and the file sample and an attention matrix between the element block sample and the task description sample; The multiple relevance labels are obtained based on at least part of the attention weights of each of the multiple attention matrices.
12. A device for determining prompt information, comprising: a relevance module, configured to determine a plurality of relevances between the task description information of the target object and a plurality of element blocks of the target file, wherein the element blocks include content blocks separated from the target file according to the file structure layout and structural layout information associated with the content blocks; a screening module, configured to screen out at least one target element block from the multiple element blocks based on the multiple relevances; The prompt information module is used to determine prompt information for the large model based on the task description information and the at least one target element block, wherein the prompt information is used to guide the large model to generate response content for the task description information.
13. An intelligent agent comprising: An input module, configured to receive input information, wherein the input information includes prompt information obtained by the method according to any one of claims 1 to 11; A processing module, configured to obtain output information by calling a large model based on the input information received by the input module; An output module is used to output the output information obtained by the processing module.
14. An electronic device comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 11.
15. A non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause the computer to execute the method according to any one of claims 1 to 11.
16. A computer program product, comprising a computer program, wherein when the computer program is executed by a processor, the method according to any one of claims 1 to 11 is implemented.