Task retrieval optimization method and device of intelligent agent, equipment, medium and product

By generating query rewriting tasks and optimizing the language model based on semi-rule reward functions, the problems of high inference cost of high-order search problems and inability to utilize labeled data in the prior art are solved, and efficient and accurate high-order search results are achieved.

CN119988435AActive Publication Date: 2025-05-13BEIJING INSTITUTE FOR GENERAL ARTIFICIAL INTELLIGENCE
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510474389.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-16
Publication Date
2025-05-13
Estimated Expiration
2045-04-16

AI Technical Summary

Technical Problem

Existing search systems cannot effectively solve the advanced search problem, especially due to the lack of real rewrite standard data, the supervised model cannot be directly trained, resulting in high inference costs and unusable labeled data.

Method used

By generating query rewriting tasks and obtaining the target loss function based on the semi-rule reward function, the preset language model is optimized and the target retrieval model is generated, thereby achieving efficient inference query rewriting.

Benefits of technology

It realizes the effect of achieving or exceeding large-scale commercial language models in inference query rewriting tasks, reduces the inference cost, can effectively utilize labeled data, and improves the efficiency and accuracy of high-order retrieval.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119988435A_ABST
    Figure CN119988435A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides an intelligent agent task retrieval optimization method which can be applied to the technical field of artificial intelligence. The task retrieval optimization method of the intelligent agent comprises the following steps: generating a query rewriting task according to original problem data of preset training data; obtaining a target loss function based on a semi-rule award function corresponding to the query rewriting task; a preset language model is optimized according to the target loss function, and a target retrieval model is generated; and performing task retrieval optimization for the update input problem through the target retrieval model. The embodiment of the invention further provides an agent task retrieval optimization device and equipment, a storage medium and a program product.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, specifically to the field of data processing technology, and more specifically to a task retrieval optimization method, device, equipment, medium and product for an intelligent agent. Background Art

[0002] Artificial Intelligence (AI) is an important driving force for the new round of scientific and technological revolution and industrial transformation. It is a new key technical science that studies and develops theories, methods, technologies and application systems for simulating, extending and expanding human intelligence. As an important part of intelligent science, artificial intelligence attempts to understand the essence of intelligence and produce a new intelligent machine (i.e., intelligent agent) that can respond in a similar way to human intelligence.

[0003] Information Retrieval (IR) is the process of finding information related to a specific query. In traditional technology, information retrieval mainly relies on keyword matching and Boolean logic, while the introduction of artificial intelligence has brought new possibilities for information retrieval. Combining information retrieval with artificial intelligence to achieve intelligent search is an important aspect of the field of artificial intelligence. Specifically, it may involve natural language processing, machine learning, deep learning and other technologies in order to significantly improve the efficiency and accuracy of information retrieval. For information retrieval in the field of artificial intelligence, existing retrieval systems are still mainly based on semantic or text matching methods, which cannot better solve high-order retrieval problems. Although query rewriting through language models can theoretically solve related problems, due to the lack of real rewriting standard data, supervised model training cannot be performed directly, making the model unable to achieve efficient and accurate high-order retrieval. Summary of the invention

[0004] In view of at least one of the above problems, embodiments of the present invention aim to provide task retrieval optimization methods, devices, equipment, media and products for intelligent agents that can achieve more intelligent, efficient, accurate and high-order information retrieval, thereby providing an open source language model of smaller size that can be trained efficiently and achieve or exceed the effects of large commercial language models in inferential query rewriting tasks.

[0005] One aspect of an embodiment of the present invention provides a task retrieval optimization method for an intelligent agent, which includes: generating a query rewriting task based on original question data of preset training data; obtaining a target loss function based on a semi-regular reward function corresponding to the query rewriting task; optimizing a preset language model according to the target loss function to generate a target retrieval model; and performing task retrieval optimization for an updated input question through the target retrieval model.

[0006] According to one embodiment of the present invention, in generating a query rewriting task based on original question data of preset training data, it includes: obtaining the original question data of the preset training data and corresponding question-related documents; generating a query rewriting task corresponding to the question-related documents according to a preset language model.

[0007] According to an embodiment of the present invention, before obtaining the target loss function based on the semi-regular reward function corresponding to the query rewriting task, it also includes: generating a semi-regular reward function according to the query rewriting task.

[0008] According to one embodiment of the present invention, in generating a semi-regular reward function based on a query rewriting task, it includes: obtaining a first correlation function generated by the query rewriting task and a second correlation function generated by original question data; generating a semi-regular reward function based on question-related documents corresponding to the original question data, the first correlation function and the second correlation function.

[0009] According to one embodiment of the present invention, in obtaining a target loss function based on a semi-regular reward function corresponding to a query rewriting task, it includes: generating a reward value corresponding to the query rewriting task through a semi-regular reward function; and estimating relative advantage information of a preset language model through the reward value.

[0010] According to an embodiment of the present invention, obtaining the target loss function based on the semi-regular reward function corresponding to the query rewriting task also includes: generating a target loss function that complies with a preset group of relative strategy optimization rules through relative advantage information.

[0011] Another aspect of the embodiments of the present invention provides a task retrieval optimization device for an intelligent agent, which includes a task generation module, a function acquisition module, a model generation module and a retrieval optimization module. The task generation module is used to generate a query rewriting task based on the original question data of the preset training data; the function acquisition module is used to obtain a target loss function based on the semi-regular reward function corresponding to the query rewriting task; the model generation module is used to optimize the preset language model according to the target loss function to generate a target retrieval model; the retrieval optimization module is used to perform task retrieval optimization for the updated input question through the target retrieval model.

[0012] Another aspect of an embodiment of the present invention provides an electronic device, including one or more processors and a memory, the memory being used to store one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors execute the above-mentioned task retrieval optimization method of the intelligent agent.

[0013] Another aspect of an embodiment of the present invention provides a computer-readable storage medium having executable instructions stored thereon, which, when executed by a processor, causes the processor to execute the above-mentioned task retrieval optimization method for an intelligent agent.

[0014] Another aspect of an embodiment of the present invention provides a computer program product, including a computer program, which implements the above-mentioned task retrieval optimization method of the intelligent agent when executed by a processor.

[0015] The task retrieval optimization method of an intelligent agent provided by an embodiment of the present invention can at least partially solve the problem that the existing retrieval system in the related art cannot better solve the high-order retrieval problem, and thus can achieve at least one of the following technical effects: Compared with the traditional information retrieval scheme, the task retrieval optimization method of the above-mentioned intelligent agent in the embodiment of the present invention can be based on the semi-regular reward function of the relevance of the mature semantic model, so that the reasoning query rewriting task for high-order information retrieval can be trained based on relevant documents through the reinforcement learning method, which solves the problem that the task lacks real rewriting standard data and cannot be directly trained with a supervised model. In addition, using the above-mentioned semi-regular reward method, the parameters of a smaller open source language model (such as Qwen2.5-1.5B) can be efficiently trained through reinforcement learning, and the computing resource requirements are low. The performance of the model after training can reach an effect close to that of a large-scale commercial language model (such as GPT-4o), realizing efficient information retrieval optimization based on a small-size language model. Moreover, through the reasoning query rewriting, the high-order retrieval task based on causal reasoning is converted into a traditional task based on text or semantic relevance. The rewritten query can be directly used for semantic retrieval or text retrieval, and the recall or ranking effect is better than the original query, that is, the effect of converting the high-order retrieval task into a conventional retrieval task is achieved.

[0016] Therefore, the task retrieval optimization method of the above-mentioned intelligent agent in the embodiment of the present invention can efficiently train a smaller open source language model, and achieve or exceed the effect of a large commercial language model in the inference query rewriting task, so as to ultimately achieve efficient and high-accuracy high-order retrieval.

[0017] It should be understood that the above general description and the following detailed description are merely exemplary and illustrative and are not intended to limit the scope of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] The above contents and other objects, features and advantages of the present invention will become more apparent through the following description of the embodiments of the present invention with reference to the accompanying drawings, in which: Figure 1 A diagram schematically illustrates an application scenario of a task retrieval optimization method, apparatus, device, medium, and program product of an intelligent agent according to an embodiment of the present invention; Figure 2 A flowchart of a task retrieval optimization method for an intelligent agent according to an embodiment of the present invention is schematically shown; Figure 3 A flowchart of an application scenario of a task retrieval optimization method for an intelligent agent according to an embodiment of the present invention is schematically shown; Figure 4 A schematic diagram showing a structural block diagram of a task retrieval optimization device for an intelligent agent according to an embodiment of the present invention; and Figure 5 A block diagram of an electronic device suitable for implementing a task retrieval optimization method for an intelligent agent according to an embodiment of the present invention is schematically shown.

[0019] The above-mentioned drawings are part of the specification of the embodiments of the present invention, which illustrate exemplary embodiments of the present invention. The attached drawings and the description of the specification are used together to illustrate the principles of the embodiments of the present invention. It should be understood that the above general description of the drawings and the following specific implementations are only exemplary and illustrative, and they cannot limit the scope of the present invention. DETAILED DESCRIPTION

[0020] In order to make the objectives, technical solutions and advantages of the embodiments of the present invention more clearly understood, the spirit of the contents disclosed by the present invention will be clearly explained with the accompanying drawings and detailed descriptions below. After understanding the embodiments of the contents of the present invention, any technician in the relevant technical field can change and modify the techniques taught by the contents of the present invention without departing from the spirit and scope of the contents of the present invention.

[0021] The exemplary embodiments and descriptions of the present invention are used to explain the present invention, but are not intended to limit the present invention. In addition, elements / components with the same or similar reference numerals used in the drawings and embodiments are used to represent the same or similar parts.

[0022] The terms “first”, “second”, etc. used in the present invention do not particularly refer to an order or sequence, nor are they used to limit the present invention. They are only used to distinguish elements or operations described with the same technical terms.

[0023] The directional terms used in the present invention, such as up, down, left, right, front or back, etc., are only used to refer to the directions of the drawings. Therefore, the directional terms used are used to illustrate and not to limit the present invention.

[0024] The words “include,” “including,” “have,” “contain,” etc. used in the present invention are open-ended terms, meaning including but not limited to.

[0025] The term "and / or" used in the present invention includes any or all combinations of the items mentioned.

[0026] Regarding the present invention, "plurality" includes "two" and "more than two"; regarding the present invention, "plurality of groups" includes "two groups" and "more than two groups".

[0027] The terms "substantially" and "approximately" used in the present invention are used to modify any quantity or error that may vary slightly, but these slight changes or errors do not change their essence. Generally speaking, the range of slight changes or errors modified by such terms may be 20% in some embodiments, 10% in some embodiments, 5% in some embodiments, or other values. Those skilled in the art should understand that the aforementioned values ​​can be adjusted according to actual needs and are not limited thereto.

[0028] All terms (including technical and scientific terms) used herein have the meanings commonly understood by those skilled in the art unless otherwise defined. It should be noted that the terms used herein should be interpreted as having a meaning consistent with the context of this specification and should not be interpreted in an idealized or overly rigid manner.

[0029] In the case of using expressions such as "at least one of A, B, and C, etc.", it should generally be interpreted in accordance with the meaning of the expression generally understood by those skilled in the art (for example, "a system having at least one of A, B, and C" should include but not be limited to a system having A alone, B alone, C alone, A and B, A and C, B and C, and / or a system having A, B, C, etc.). In the case of using expressions such as "at least one of A, B, or C, etc.", it should generally be interpreted in accordance with the meaning of the expression generally understood by those skilled in the art (for example, "a system having at least one of A, B, or C" should include but not be limited to a system having A alone, B alone, C alone, A and B, A and C, B and C, and / or a system having A, B, C, etc.). Those skilled in the art should also understand that any transitional conjunctions and / or phrases that substantially represent two or more optional items, whether in the specification, claims, or drawings, should be understood to give the possibility of including one of these items, either of these items, or both of these items. For example, the phrase "A or B" should be understood to include the possibilities of "A" or "B", or "A and B".

[0030] Relevance is a core issue in the field of information retrieval. Given a query, the goal of the retrieval system is to recall documents related to the query from a large number of documents, sort them by relevance, and return them to the user. Traditional information retrieval systems mainly focus on relevance at the text or semantic level, but in recent years, with the development and application of new forms of products such as language models and deep research (such as DeepResearch), information retrieval systems need to retrieve effective answers based on causal relevance for given questions - there may not be a significant correlation between questions and answers at the text or semantic level. This type of problem can be called a "high-order retrieval problem". Existing retrieval systems are mainly based on semantic or text matching methods and cannot solve this type of high-order retrieval problem well.

[0031] Specifically, for level 1, conventional retrieval tasks are usually implemented based on keyword retrieval. For example, if the retrieval question is "What is the widest highway in North America?", the positive document "The part of highway 401 that passes through Toronto is North America's busiest highway, and one of the widest" can be obtained through the relevance judgment of keyword matching. Furthermore, for level 2, retrieval tasks are usually implemented based on semantic retrieval. For example, if the retrieval question is "How human activities influence climate system?", the positive document "Deforestation and urbanization result in increased emissions, urban heat island effects and changes in natural watercycle" can be obtained through the relevance judgment of semantic matching.

[0032] For high-order retrieval questions (level 3), they can be implemented based on reasoning-based retrieval. For example, retrieval questions can include sustainable living-post, code-issue and math-question. For example, sustainable living-post is “At home, after l water my plants, thewater goes to plates below the pots. Can l rouse it for my plants next time”; code-issue is “I have this table and need to transform it to ... l don't like UNPIVOT. Is there a better function in snowflake for this”; math-question is “Let k=2008^2+2^2008. What is the units digit of k*2+2^k?”. Then for the Sustainable living-post relevance “Risk of using recycled plant water”, the positive document is “Soluble salts are commonly found insoils When they build up, they destroy the soil structure and cause direct damage to roots”; for the Code-issue relevance “Alternative function”, the positive document is “The function FLATTEN flattens (explodes) compound values ​​into multiple rows... FLATTEN INPUT→ <expr>..."; for the Math-question relevance "Uses the same theorem", the corresponding positive sample document is "Determine all positive integers relatively prime to all the terms of the infinite sequence a_n=2^n+3^n+6^n -1...". Among them, the positive sample document can be understood as the correct answer to the retrieval question.

[0033] A feasible solution is to use the causal reasoning ability of the language model to perform reasoning rewrite on the input question and generate a rewritten new query - usually the query is closer to the relevant documents at the text or semantic level. Therefore, using the rewritten query, you can directly recall relevant documents based on the existing text or semantic retrieval system to achieve better retrieval results without changing the entire retrieval system framework. However, this type of query rewriting method relies on a language model with a large number of parameters (such as GPT-4o or Llama3-70B), which has a high inference cost and is difficult to deploy in practice. A more ideal method is to train a query rewriting model with a small number of parameters based on a small amount of annotated data, but because manual data annotation can only realize the correlation annotation between questions and documents, it is impossible to directly write a rewritten query. Therefore, traditional supervised learning methods cannot be directly used for this task.

[0034] Therefore, this results in at least the following technical problems in task retrieval solutions for high-level retrieval problems: (1) Excessive inference cost: Existing causal reasoning query rewriting methods directly rely on high-performance language models, which have high inference costs and are difficult to improve in rewriting effects.

[0035] (2) Inability to effectively utilize labeled data: For retrieval tasks, the labeled data is in the form of relevance labels between the original query and the document (e.g., relevant or irrelevant). Such labeled data can be used to train discriminative models (e.g., recall or ranking models), but cannot be used to train generative models. Therefore, even if the current methods have labeled data and training resources, they cannot be used for model training.

[0036] In view of at least one of the above problems, embodiments of the present invention aim to provide task retrieval optimization methods, devices, equipment, media and products for intelligent agents that can achieve more intelligent, efficient, accurate and high-order information retrieval, thereby providing an open source language model of smaller size that can be trained efficiently and achieve or exceed the effects of large commercial language models in inferential query rewriting tasks.

[0037] One aspect of an embodiment of the present invention provides a task retrieval optimization method for an intelligent agent, which includes: generating a query rewriting task based on original question data of preset training data; obtaining a target loss function based on a semi-regular reward function corresponding to the query rewriting task; optimizing a preset language model according to the target loss function to generate a target retrieval model; and performing task retrieval optimization for an updated input question through the target retrieval model.

[0038] Figure 1 An application scenario diagram of a task retrieval optimization method, apparatus, device, medium, and program product of an intelligent agent according to an embodiment of the present invention is schematically shown.

[0039] like Figure 1 As shown, the application scenario 100 according to this embodiment may include terminal devices 101, 102, 103, a network 104 and a server 105. The network 104 is used to provide a medium for a communication link between the terminal devices 101, 102, 103 and the server 105. The network 104 may include various connection types, such as wired, wireless communication links or optical fiber cables, etc.

[0040] Users can use terminal devices 101, 102, 103 to interact with server 105 through network 104 to receive or send messages, etc. Various communication client applications can be installed on terminal devices 101, 102, 103, such as shopping applications, web browser applications, search applications, instant messaging tools, email clients, social platform software, etc. (only examples).

[0041] The terminal devices 101 , 102 , and 103 may be various electronic devices having a display screen and supporting web browsing, including but not limited to smart phones, tablet computers, laptop computers, and desktop computers.

[0042] The server 105 may be a server that provides various services, such as a background management server (only as an example) that provides support for websites browsed by users using the terminal devices 101, 102, and 103. The background management server may analyze and process the received data such as user requests, and feed back the processing results (such as web pages, information, or data obtained or generated according to user requests) to the terminal device.

[0043] It should be noted that the task retrieval optimization method of the intelligent agent provided in the embodiment of the present invention can generally be executed by the server 105. Accordingly, the task retrieval optimization device of the intelligent agent provided in the embodiment of the present invention can generally be set in the server 105. The task retrieval optimization method of the intelligent agent provided in the embodiment of the present invention can also be executed by a server or server cluster that is different from the server 105 and can communicate with the terminal devices 101, 102, 103 and / or the server 105. Accordingly, the task retrieval optimization device of the intelligent agent provided in the embodiment of the present invention can also be set in a server or server cluster that is different from the server 105 and can communicate with the terminal devices 101, 102, 103 and / or the server 105.

[0044] It should be understood that Figure 1 The number of terminal devices, networks and servers in the embodiment is only for illustration. Any number of terminal devices, networks and servers may be provided according to implementation requirements.

[0045] The following will be based on Figure 1 The scene described by Figure 2~Figure 3 The task retrieval optimization method of the intelligent agent in the disclosed embodiment is described in detail.

[0046] like Figure 2 As shown, one aspect of an embodiment of the present invention provides a task retrieval optimization method for an intelligent agent, which includes operations S201 to S204.

[0047] In operation S201, a query rewriting task is generated according to original question data of preset training data; In operation S202, a target loss function is obtained based on a semi-regular reward function corresponding to the query rewriting task; In operation S203, the preset language model is optimized according to the target loss function to generate a target retrieval model; and In operation S204 , task retrieval optimization is performed on the updated input question through the target retrieval model.

[0048] The intelligent agent can be the executor of the above-mentioned task retrieval optimization method in the embodiment of the present invention, or it can be an executor controlled by the task retrieval optimization method. Specifically, it can be a humanoid intelligent robot or other AI device, which usually has its own executor to complete specific action tasks.

[0049] For high-order retrieval tasks, it is necessary to build a targeted retrieval model for high-order retrieval problems in order to achieve the expected efficient and accurate high-order retrieval effect. The preset training data is data trained for the retrieval model so that the retrieval model can have the task retrieval optimization capability of the embodiment of the present invention. Specifically, it can be acquired in advance through historical document data or historical experimental data (such as data capture and preprocessing), or it can be manually set in advance. Among them, the high-order retrieval task can be understood as an intelligent retrieval task for high-order retrieval problems. The high-order retrieval task can specifically be performed through the retrieval model generated after the preset training data training optimization to perform retrieval processing, and finally obtain the corresponding retrieval results.

[0050] The preset training data may generally include the original question data and the question-related documents corresponding to the original question data. Among them, the original question data can be understood as various multimodal data such as retrieval text (such as characters, words, sentences and paragraphs), images, audio and video for question-related documents, which can be understood as the retrieval question of the high-order retrieval task, that is, Question. In addition, the question-related documents can correspond to the standard retrieval answer of the above original question data, that is, Answer, and the question-related documents can also be embodied in various multimodal data such as text (such as characters, words, sentences and paragraphs), images, audio and video. It should be noted that the preset training data composed of the original question data and the question-related documents can be ordinary retrieval training data in conventional information retrieval technology.

[0051] The query rewriting task can be a query task after rewriting the original question based on semantic reasoning (Reasoning-Based) or even image reasoning, i.e. Querry. Specifically, it can be obtained by rewriting the original question data of the preset training data, wherein the so-called rewriting process can usually be based on the preset language model through the semantic reasoning process to realize the semantic extraction of the original question data and generate the query rewriting task, and after the query rewriting, the query rewriting task corresponds to the question-related document corresponding to the original question data, which can be understood as the query rewriting task generated by the original question data after the query rewriting and the corresponding question-related document have a corresponding relationship.

[0052] The semi-regular reward function may be a reward function related to the relevance of the retrieval result prediction for the query rewriting task and the original question data of the preset training data. Specifically, the retrieval result prediction may be implemented by a mature semantic model. The target loss function may be a loss function (Loss Function) optimized for training the preset language model.

[0053] The preset language model can be a relatively small open source language model (such as Qwen2.5-1.5B) implemented by reinforcement learning. The preset language model requires a relatively low amount of computing resources. Before the optimization process of the above-mentioned operation S203 of the embodiment of the present invention, its data processing performance is usually not comparable to that of a language model such as GPT-4o. The preset language model is optimized by the above-mentioned target loss function, and a target retrieval model can be finally generated. The target retrieval model can specifically be a retrieval model provided by the task retrieval optimization method in the embodiment of the present invention to achieve high-order retrieval, and can specifically be obtained by optimizing the preset language model through the target loss function. Therefore, it is possible to achieve efficient training of a relatively small open source language model (such as Qwen2.5-1.5B) through ordinary retrieval problem data to ensure that the target retrieval model can reach or exceed the retrieval performance of a large commercial language model (such as GPT-4o or Llama3-70B) in the inference query rewriting task.

[0054] The updated input question can be one or more original question data in the original question data set of the preset training data (which can also be used as the verification data of the target retrieval model), or other retrieval question data different from the preset training data. The updated input question is processed by the target retrieval model, and finally the retrieval capability of the large language model can be achieved at least in the inference query rewriting task based on the updated input question, and the best retrieval results can be obtained, so as to achieve the purpose of task retrieval optimization.

[0055] Therefore, the task retrieval optimization method of the above-mentioned intelligent agent in the embodiment of the present invention can get rid of the dependence on high-performance and high-cost language models (such as GPT-4o or Llama3-70B), the reasoning cost is lower, the query rewriting effect can be greatly improved, and it is easier to deploy in actual application scenarios. In addition, the target retrieval model can perform efficient model training based on the query rewriting of the original question data, and realize the reinforcement learning process based on semi-rule rewards, thereby solving the problem that the labeled data in the inference rewriting task in the traditional technical solution cannot be used for model training. Finally, it is possible to efficiently train a smaller open source language model, and achieve or exceed the retrieval performance of a large commercial language model in the inference query rewriting task, so as to ultimately achieve efficient and high-accuracy high-order retrieval, and complete the optimization of the high-order retrieval process.

[0056] In order to enable those skilled in the art to have a clearer understanding of the task retrieval optimization method of the above-mentioned intelligent agent in the embodiment of the present invention, the following is further provided: Figure 3-Figure 4 Description.

[0057] like Figure 2 and Figure 3 As shown, according to an embodiment of the present invention, in operation S201, generating a query rewriting task according to original question data of preset training data includes: Obtain the original question data of the preset training data and the corresponding question-related documents; Generate query rewriting tasks corresponding to question-related documents based on the preset language model.

[0058] like Figure 3 As shown, for a high-level retrieval task, there may generally be a set of preset training data in the form of "original question data and corresponding question-related documents", and each corresponding "original question data 311 and corresponding question-related document 312" may be used as a preset training data 301. The original question data 311 may be text, image, audio or even video data, and the corresponding question-related document 312 may also be text, image, audio or even video data.

[0059] The original question data 311 may be a search question (i.e., Question), and the question-related document 312 may be a standard answer (i.e., Answer) corresponding to the original question data, which may be specifically represented as: ,in is a positive integer greater than or equal to 1. For example, for a level 1 conventional retrieval task, the original question data may be "What is the widest highway in North America?", and the corresponding question-related document may be "The part of highway 401 that passes through Toronto is North America's busiest highway, and one of the widest", while for a level 3 high-level retrieval task, the corresponding original question data and question-related documents may be designed based on the aforementioned Sustainable living-post, Code-issue, and Math-question, which will not be described in detail.

[0060] The preset language model can be used as a language model (such as Qwen2.5-1.5B) for query rewriting for the original question data. Specifically, the query rewriting task can be expressed as the following formula 1:

[0061] Wherein, query is the query rewriting task 302 after rewriting (such as Figure 3 As shown in ), LLM is the preset language model for query rewriting, and question is the original question data 311 (as shown in Figure 3 as shown).

[0062] After the above preset language model LLM rewriting, a mapping relationship between the query rewriting task query and the corresponding question-related document doc can be established, so that the query rewriting task query can directly recall the corresponding question-related document through the existing text or semantic retrieval recall / ranking system , and the recall effect can be significant due to the retrieval method that directly uses the original question data for recall.

[0063] Furthermore, according to the generation process of the query rewriting task, for each batch of preset training data, the preset language model to be trained is used to perform multiple rounds of data generation on the original question data question in the input preset training data, so that multiple different query rewriting tasks query can be repeatedly generated.

[0064] Therefore, by presetting the language model, efficient query rewriting can be achieved for routine original question data, and the generation of high-order retrieval tasks can be achieved. Specifically, through inference-based query rewriting, high-order retrieval tasks based on causal reasoning are converted into traditional tasks based on text or semantic relevance. The rewritten query can be directly used for semantic retrieval or text retrieval, achieving recall or ranking effects that are better than the original query.

[0065] like Figure 2 and Figure 3 As shown, according to an embodiment of the present invention, before operation S202 obtains the target loss function based on the semi-regular reward function corresponding to the query rewriting task, it also includes: Generate a semi-regular reward function based on the query rewriting task.

[0066] In the embodiment of the present invention, the semi-regular reward can be understood as the calculation of the relevance score of the query rewriting task in the reward function, which can be implemented by the relevance model, rather than a purely rule-based reward function. Based on the above semi-regular reward function, a reinforcement learning method is used to train a preset open source language model of a smaller size, so that a reasoning query rewriting model with better performance, namely, a target retrieval model, can be obtained.

[0067] like Figure 2 and Figure 3 As shown, according to an embodiment of the present invention, in generating a semi-regular reward function according to a query rewriting task, the following steps are included: Obtaining a first correlation function generated by the query rewriting task and a second correlation function generated by the original question data; A semi-regular reward function is generated based on question-related documents corresponding to the original question data, a first relevance function, and a second relevance function.

[0068] In one embodiment of the present invention, a query rewriting process such as the above formula 1 can be performed on each input original question data to generate a corresponding query rewriting task. Figure 3 As shown, further, for the original question data 311, a corresponding second correlation function 304 can be obtained through a mature semantic model (such as the BGE series model) or a heuristic correlation algorithm (such as BM25); correspondingly, for the query rewriting task 302, the corresponding first correlation function 303 can be generated through the same semantic model or heuristic algorithm.

[0069] Therefore, combined with the query rewriting task query corresponding to all positive sample question-related documents covering the real user intention , the corresponding semi-regular reward function 305 (such as Figure 3 As shown in the figure, it is expressed as the following formula 2:

[0070] in, It can be understood as a reward score. The set of all positive sample documents that cover the real user intent corresponding to the query rewriting task query can be used. It can be a correlation calculation function. The correlation calculation function can be implemented by a mature semantic model (such as the BGE series model) or by a heuristic correlation algorithm (such as BM25). Therefore, a semi-regular reward construction method based on semantic correlation can be implemented to perform reinforcement learning training on the language model and inject positive sample information into the model by an unsupervised learning method.

[0071] The calculation logic of the semi-regular reward function of Formula 2 above can be expressed as: The query rewriting task generated by the inference rewriting , the text and semantic level should have more prominent intention information, which is more easily captured by the relevance model (relevance calculation function ) Capture - For each positive sample question related documents , query rewrite task The relevance score should be as high as possible from the original question data Therefore, for the relevance score of documents related to the positive question, the query rewriting task Relative to the original problem data The added value is the query rewrite task The reward score obtained. The relevance may be a score between 0 and 1.

[0072] Here, "semi-rule reward" refers to the calculation of the relevance score in the reward function in the above formula 2, which can specifically rely on another relevance model, not a purely rule-based reward function. It should be noted that, compared with the process reward model commonly used in language model reinforcement learning training, the mature retrieval relevance model used in the embodiment of the present invention also has the advantages of low computational cost and good robustness of the rule reward function. In addition, the rewriting result (i.e., the query rewriting task) used in the calculation of the reward ) is the final output of the model, which does not involve any intermediate process (such as the steps in solving a math problem), and is also immune to reward hacking like a rule reward function. Therefore, it can be defined as a "semi-rule reward" here. To facilitate calculation and improve robustness, a semantic-based relevance model can be introduced. Among them, if the calculation of the relevance score depends on a heuristic algorithm (such as BM25), then the reward function can also become a completely rule-based reward function.

[0073] Therefore, by constructing a semi-regular reward method based on semantic relevance (i.e., the relevance semi-regular reward function of a mature semantic model), the inference-based query rewriting task for high-order information retrieval can be implemented through reinforcement learning and model training based on relevant documents, thereby solving the problem that the task lacks real rewriting annotated data and cannot directly perform supervised model training.

[0074] like Figure 2 and Figure 3 As shown, according to an embodiment of the present invention, in operation S202, obtaining a target loss function based on a semi-regular reward function corresponding to a query rewriting task includes: Generate the reward value corresponding to the query rewriting task through a semi-regular reward function; The relative advantage information of the preset language model is estimated through the reward value.

[0075] In one embodiment of the present invention, the semi-regular reward function provided by the above formula 2 can calculate the reward values ​​of all query rewriting task queries corresponding to the preset training data one by one, and estimate the relative advantage information (Advantage) of the retrieval action corresponding to the query rewriting task calculated by the preset language model through the reward values ​​of all query rewriting task queries. The relative advantage information can be understood as the corresponding reward score (i.e., Reward Score) of the above semi-reward rule function.

[0076] like Figure 2 and Figure 3 As shown, according to an embodiment of the present invention, in operation S202, obtaining the target loss function based on the semi-regular reward function corresponding to the query rewriting task also includes: The target loss function that conforms to the preset group relative strategy optimization rules is generated through relative advantage information.

[0077] Furthermore, if Figure 3 As shown, the above relative advantage information can be used to calculate the target loss function 306 based on the preset group relative strategy optimization rule, thereby optimizing the preset language model and ultimately generating the target retrieval model through the target loss function.

[0078] The preset group relative policy optimization rule may specifically be a reinforcement learning algorithm rule based on group relative policy optimization (GroupRelative Policy Optimization, GRPO for short), which may be used to enhance the reasoning capability of the preset language model.

[0079] Therefore, by using a reinforcement learning method based on a preset group relative policy optimization rule and a semi-regular reward function, the preset language model used for query rewriting can be trained and optimized to generate a target retrieval model.

[0080] This can further ensure the implementation of a semi-regular reward construction method based on semantic relevance, which can be used to perform reinforcement learning training on the language model and inject positive sample information into the model using an unsupervised learning method.

[0081] In summary, based on the above semi-regular reward function, the target model of inference query rewriting with better performance can be obtained by training the smaller open source language model using the GRPO-based reinforcement learning method. Specifically, through this semi-regular reward function, efficient parameter training can be achieved for smaller open source language models (such as Qwen2.5-1.5B) through reinforcement learning, with lower computing resource requirements, and the performance of the trained model can be close to that of large-scale commercial language models (such as GPT-4o).

[0082] Compared with the traditional information retrieval scheme, the task retrieval optimization method of the above-mentioned intelligent agent in the embodiment of the present invention can be based on the semi-regular reward function of the relevance of the mature semantic model, so that the reasoning query rewriting task for high-order information retrieval can be trained based on relevant documents through the reinforcement learning method, which solves the problem that the task lacks real rewriting standard data and cannot be directly trained with a supervised model. In addition, using the above-mentioned semi-regular reward method, the parameters of a smaller open source language model (such as Qwen2.5-1.5B) can be efficiently trained through reinforcement learning, and the computing resource requirements are low. The performance of the model after training can reach an effect close to that of a large-scale commercial language model (such as GPT-4o), realizing efficient information retrieval optimization based on a small-size language model. Moreover, through the reasoning query rewriting, the high-order retrieval task based on causal reasoning is converted into a traditional task based on text or semantic relevance. The rewritten query can be directly used for semantic retrieval or text retrieval, and the recall or ranking effect is better than the original query, that is, the effect of converting the high-order retrieval task into a conventional retrieval task is achieved.

[0083] Therefore, the task retrieval optimization method of the above-mentioned intelligent agent in the embodiment of the present invention can efficiently train a smaller open source language model, and achieve or exceed the effect of a large commercial language model in the inference query rewriting task, so as to ultimately achieve efficient and high-accuracy high-order retrieval.

[0084] In the task retrieval optimization method of the above-mentioned intelligent agent in the embodiment of the present invention, the core technical points are the construction and application of the semi-regular reward function and the model training process for inference retrieval. Specifically, a semi-regular reward construction method based on semantic relevance is used to perform reinforcement learning training on the language model, and the positive sample information is injected into the model by an unsupervised learning method. In addition, the model training process is used to enable the model to have the ability of inference query rewriting through reinforcement learning. Therefore, it is possible to efficiently train a smaller open source language model to achieve or exceed the effect of a large commercial language model on the task of inference query rewriting.

[0085] Based on the above-mentioned task retrieval optimization method of an intelligent agent, the present invention also provides a task retrieval optimization device for an intelligent agent. Figure 4 The device 400 is described in detail.

[0086] Figure 4 The structural block diagram of the task retrieval optimization device 400 of an intelligent agent according to an embodiment of the present invention is schematically shown.

[0087] like Figure 4 As shown, the task retrieval optimization device 400 of the intelligent agent in this embodiment includes a task generation module 410, a function acquisition module 420, a model generation module 430 and a retrieval optimization module 440.

[0088] The task generation module 410 is used to generate a query rewriting task according to the original question data of the preset training data. In one embodiment, the task generation module 410 can be used to perform the operation S201 described above, which will not be described in detail here.

[0089] The function acquisition module 420 is used to acquire the target loss function based on the semi-regular reward function corresponding to the query rewriting task. In one embodiment, the function acquisition module 420 can be used to perform the operation S202 described above, which will not be repeated here.

[0090] The model generation module 430 is used to optimize the preset language model according to the target loss function to generate a target retrieval model. In one embodiment, the model generation module 430 can be used to perform the operation S203 described above, which will not be described in detail here.

[0091] The retrieval optimization module 440 is used to perform task retrieval optimization for the updated input question through the target retrieval model. In one embodiment, the retrieval optimization module 440 can be used to perform the operation S204 described above, which will not be described in detail here.

[0092] According to an embodiment of the present invention, any multiple modules among the task generation module 410, the function acquisition module 420, the model generation module 430 and the retrieval optimization module 440 can be combined into one module for implementation, or any one of the modules can be split into multiple modules. Alternatively, at least part of the functions of one or more of these modules can be combined with at least part of the functions of other modules and implemented in one module. According to an embodiment of the present invention, at least one of the task generation module 410, the function acquisition module 420, the model generation module 430 and the retrieval optimization module 440 can be at least partially implemented as a hardware circuit, such as a field programmable gate array (FPGA), a programmable logic array (PLA), a system on a chip, a system on a substrate, a system on a package, an application specific integrated circuit (ASIC), or can be implemented by hardware or firmware such as any other reasonable way of integrating or packaging the circuit, or implemented in any one of the three implementation methods of software, hardware and firmware or in a proper combination of any of them. Alternatively, at least one of the task generation module 410 , the function acquisition module 420 , the model generation module 430 , and the retrieval optimization module 440 may be at least partially implemented as a computer program module, which may perform a corresponding function when executed.

[0093] Figure 5 A block diagram of an electronic device suitable for implementing a task retrieval optimization method for an intelligent agent according to an embodiment of the present invention is schematically shown.

[0094] The electronic device provided by an embodiment of the present invention includes one or more processors and a memory, the memory being used to store one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors execute the task retrieval optimization method of the above-mentioned intelligent agent.

[0095] like Figure 5 As shown, the electronic device 500 according to an embodiment of the present invention includes a processor 501, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 502 or a program loaded from a storage part 508 to a random access memory (RAM) 503. The processor 501 may include, for example, a general-purpose microprocessor (such as a CPU), an instruction set processor and / or a related chipset and / or a special-purpose microprocessor (for example, an application-specific integrated circuit (ASIC)), etc. The processor 501 may also include an onboard memory for caching purposes. The processor 501 may include a single processing unit or multiple processing units for performing different actions of the method flow according to an embodiment of the present invention.

[0096] In RAM 503, various programs and data required for the operation of electronic device 500 are stored. Processor 501, ROM 502 and RAM 503 are connected to each other via bus 504. Processor 501 performs various operations of the method flow according to the embodiment of the present invention by executing the program in ROM 502 and / or RAM 503. It should be noted that the program can also be stored in one or more memories other than ROM 502 and RAM 503. Processor 501 can also perform various operations of the method flow according to the embodiment of the present invention by executing the program stored in the one or more memories.

[0097] According to an embodiment of the present invention, the electronic device 500 may further include an input / output (I / O) interface 505, which is also connected to the bus 504. The electronic device 500 may further include one or more of the following components connected to the I / O interface 505: an input portion 506 including a keyboard, a mouse, etc.; an output portion 507 including a cathode ray tube (CRT), a liquid crystal display (LCD), etc., and a speaker, etc.; a storage portion 508 including a hard disk, etc.; and a communication portion 509 including a network interface card such as a LAN card, a modem, etc. The communication portion 509 performs communication processing via a network such as the Internet. A drive 510 is also connected to the I / O interface 505 as needed. A removable medium 511, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 510 as needed, so that a computer program read therefrom is installed into the storage portion 508 as needed.

[0098] The present invention also provides a computer-readable storage medium on which executable instructions are stored. When the instructions are executed by a processor, the processor executes the task retrieval optimization method of the above-mentioned intelligent agent.

[0099] The computer-readable storage medium may be included in the device / apparatus / system described in the above embodiment; or it may exist independently without being assembled into the device / apparatus / system. The above computer-readable storage medium carries one or more programs, and when the above one or more programs are executed, the method according to the embodiment of the present invention is implemented.

[0100] According to an embodiment of the present invention, the computer-readable storage medium may be a non-volatile computer-readable storage medium, for example, may include but is not limited to: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In the present invention, the computer-readable storage medium may be any tangible medium containing or storing a program, which may be used by or in combination with an instruction execution system, an apparatus or a device. For example, according to an embodiment of the present invention, the computer-readable storage medium may include the ROM 502 and / or RAM 503 described above and / or one or more memories other than ROM 502 and RAM 503.

[0101] An embodiment of the present invention also includes a computer program product, which includes a computer program, and when the computer program is executed by a processor, it implements the task retrieval optimization method of the above-mentioned intelligent agent.

[0102] The computer program includes program codes for executing the method shown in the flow chart. When the computer program product is run in a computer system, the program codes are used to enable the computer system to implement the method provided by the embodiment of the present invention.

[0103] The computer program executes the above functions defined in the system / device of the embodiment of the present invention when it is executed by the processor 501. According to the embodiment of the present invention, the system, device, module, unit, etc. described above can be implemented by a computer program module.

[0104] In one embodiment, the computer program may rely on tangible storage media such as optical storage devices, magnetic storage devices, etc. In another embodiment, the computer program may also be transmitted and distributed in the form of signals on a network medium, and downloaded and installed through the communication part 509, and / or installed from the removable medium 511. The program code contained in the computer program may be transmitted using any appropriate network medium, including but not limited to: wireless, wired, etc., or any suitable combination of the above.

[0105] In such an embodiment, the computer program can be downloaded and installed from the network through the communication part 509, and / or installed from the removable medium 511. When the computer program is executed by the processor 501, the above functions defined in the system of the embodiment of the present invention are performed. According to the embodiment of the present invention, the system, device, means, module, unit, etc. described above can be implemented by a computer program module.

[0106] According to an embodiment of the present invention, the program code for executing the computer program provided by the embodiment of the present invention can be written in any combination of one or more programming languages, specifically, these computing programs can be implemented using high-level process and / or object-oriented programming languages, and / or assembly / machine languages. Programming languages ​​include, but are not limited to, such as Java, C++, python, "C" language or similar programming languages. The program code can be executed completely on the user computing device, partially on the user device, partially on the remote computing device, or completely on the remote computing device or server. In the case of a remote computing device, the remote computing device can be connected to the user computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computing device (e.g., using an Internet service provider to connect through the Internet).

[0107] The flow chart and block diagram in the accompanying drawings illustrate the possible architecture, function and operation of the system, method and computer program product according to various embodiments of the present invention. In this regard, each box in the flow chart or block diagram can represent a module, a program segment, or a part of a code, and the above-mentioned module, program segment, or a part of a code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order from the order marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram or flow chart, and the combination of the boxes in the block diagram or flow chart can be implemented with a dedicated hardware-based system that performs a specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.

[0108] In addition, all actions of acquiring information, signals or data in the present invention are carried out in compliance with the corresponding data protection laws, regulations and policies of the country where they are located, and with the authorization given by the owner of the corresponding device.

[0109] It will be appreciated by those skilled in the art that the features described in the various embodiments and / or claims of the present invention may be combined and / or combined in various ways, even if such combinations and / or combinations are not explicitly described in the present invention. In particular, the features described in the various embodiments and / or claims of the present invention may be combined and / or combined in various ways without departing from the spirit and teachings of the present invention. All of these combinations and / or combinations fall within the scope of the present invention.

[0110] The embodiments of the present invention are described above. However, these embodiments are only for the purpose of illustration, and are not intended to limit the scope of the present invention. Although each embodiment is described above, it does not mean that the measures in each embodiment cannot be used in combination. The scope of the present invention is defined by the attached claims and their equivalents. Without departing from the scope of the present invention, those skilled in the art may make various substitutions and modifications, which should all fall within the scope of the present invention.< / expr>

Claims

1. A task retrieval optimization method for an intelligent agent, characterized in that: include: Generate query rewriting tasks based on the original question data of preset training data; Obtaining a target loss function based on a semi-regular reward function corresponding to the query rewriting task; Optimizing the preset language model according to the target loss function to generate a target retrieval model; and Task retrieval optimization is performed for the updated input problem through the target retrieval model.

2. The method according to claim 1, characterized in that: In generating the query rewriting task according to the original question data of the preset training data, it includes: Obtaining original question data and corresponding question-related documents of the preset training data; A query rewriting task corresponding to the document related to the question is generated according to the preset language model.

3. The method according to claim 1, characterized in that Before obtaining the target loss function based on the semi-regular reward function corresponding to the query rewriting task, the method further includes: The semi-regular reward function is generated according to the query rewriting task.

4. The method according to claim 3, characterized in that In generating the semi-regular reward function according to the query rewriting task, comprising: Obtaining a first correlation function generated by the query rewriting task and a second correlation function generated by the original question data; The semi-regular reward function is generated based on the question-related documents corresponding to the original question data, the first relevance function and the second relevance function.

5. The method according to claim 1, characterized in that In the obtaining of the target loss function based on the semi-regular reward function corresponding to the query rewriting task, it includes: Generating a reward value corresponding to the query rewriting task by using the semi-regular reward function; The relative advantage information of the preset language model is estimated by the reward value.

6. The method according to claim 5, characterized in that In the step of obtaining a target loss function based on a semi-regular reward function corresponding to the query rewriting task, the step further includes: The relative advantage information is used to generate a target loss function that conforms to a preset group relative strategy optimization rule.

7. A task retrieval optimization device for an intelligent agent, characterized in that: include: A task generation module, used to generate query rewriting tasks based on original question data of preset training data; A function acquisition module, used for acquiring a target loss function based on a semi-regular reward function corresponding to the query rewriting task; A model generation module, used to optimize the preset language model according to the target loss function to generate a target retrieval model; and The retrieval optimization module is used to perform task retrieval optimization for the updated input problem through the target retrieval model.

8. An electronic device comprising: one or more processors; a memory for storing one or more programs, When the one or more programs are executed by the one or more processors, the one or more processors execute the method according to any one of claims 1 to 6.

9. A computer-readable storage medium having executable instructions stored thereon, which, when executed by a processor, causes the processor to execute the method according to any one of claims 1 to 6.

10. A computer program product, comprising a computer program, wherein when the computer program is executed by a processor, the method according to any one of claims 1 to 6 is implemented.

Citation Information

Patent Citations

  • Multi-round text-to-SQL method and system based on conversation rewriting model

    CN112905637A

  • Automatic reward function design system and method based on large language model

    CN118036675A

  • Intelligent patent retrieval method and system based on Agent

    CN118760761A

  • Complex query response generation system and method based on retrieval enhancement

    CN118820295A

  • Intelligent search engine construction method based on large language model

    CN118964589A