Query expansion method based on generative retrieval
By introducing large language models into the query extension method for deep semantic understanding and keyword extraction, and appropriately weakening of query conditions during the generation process, the problems of insufficient query extension correlation and poor handling of restriction conditions in the prior art are solved, and higher relevance and recall rates of search results are achieved.
Patent Information
- Application Number
- CN202510110465.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-23
- Publication Date
- 2025-05-23
AI Technical Summary
When existing query extension methods handle complex queries, it is difficult to effectively understand the query context and global semantics, resulting in insufficient correlation of generated extended words, which may introduce unrelated content, and fail to adequately handle restrictions, resulting in low accuracy and recall of search results.
A two-stage query extension method based on generative retrieval is adopted: the first stage is to deeply analyze the query through large language models and extract core keywords; the second stage is to guide the generation process through these keywords, appropriately delete or weaken excessively strict query conditions, and replace precise terms with a wider range of terms, thereby improving the flexibility and coverage of the query.
It effectively improves the relevance and recall rate of query expansion, allowing complex queries to obtain more accurate and comprehensive search results, avoiding the 'zero result' problem caused by excessive restrictions, while maintaining accurate expression of the user's original intentions.
Smart Images

Figure CN120030137A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of information retrieval and natural language processing, and in particular relates to a query expansion method based on generative retrieval. Background Art
[0002] In information retrieval systems, when a user's query is too specific or contains too many restrictions, it often results in "zero results" or less relevant search results. To solve this problem, query rewriting technology optimizes the search effect by adjusting the expression of the query, among which query expansion is a commonly used method. Existing query expansion methods mainly enhance the semantic coverage of the query by adding synonyms or context-related words, thereby improving the recall rate. However, these methods have the following defects in practical applications:
[0003] 1. Insufficient relevance and contextual understanding of expanded terms: Existing query expansion methods usually rely on static dictionaries, statistical models or simple context analysis. The generated expanded terms are often weakly related to the actual needs of the query subject and may introduce irrelevant content. More importantly, these methods lack a comprehensive understanding of the query context and global semantics, resulting in the generated expanded terms failing to fully reflect the user's true intentions, affecting the precision and recall of the retrieval results.
[0004] (1) Existing query expansion methods may generate overly broad expansion terms.
[0005] For example, when querying "high performance computing and data storage", traditional query expansion methods may generate words such as "computer" or "database", which are relevant to the query in some cases, but they are too general and may introduce a large number of irrelevant search results, resulting in a poor user experience. Users may expect more precise results, such as content about specific storage technologies (such as "distributed storage" or "solid state drives"), but traditional methods fail to fully understand this, resulting in insufficient relevance of the expanded terms.
[0006] (2) Existing query expansion methods fail to fully understand the semantic associations of multiple key components in a query.
[0007] For example, in a medical search scenario, when a user enters "prevention and drug treatment of acute asthma in children", traditional expansion methods often fail to fully understand the semantic associations of multiple key components in the query, such as "children", "acute asthma" and "drug treatment". This lack of contextual understanding may cause the system-generated expansion terms (such as "asthma" or "treatment") to be too broad and fail to accurately identify specific treatment plans or drugs for children, ultimately affecting the quality of the search results.
[0008] 2. Insufficient handling of restrictions: Too many restrictions in complex queries may significantly narrow the search scope. Existing methods mainly focus on expanding the semantic coverage of queries, but fail to effectively identify and alleviate the restrictions on the search scope imposed by these conditions. Conditions such as time, domain, and type may lead to very limited search results. Even if some semantic expressions are expanded, it is still difficult to significantly improve the recall rate.
[0009] (1) The user’s true intention may contain a certain degree of implicit flexibility, and the user’s expectations for search results may be broad.
[0010] For example, in the legal field, users may search for "administrative penalty cases on environmental pollution in 2023". When there are fewer penalty cases in 2023 or no public data, the system can appropriately relax the time limit and provide relevant cases at the end of 2022 or the beginning of 2024 to help users obtain more information. Entering "2023" does not necessarily mean completely excluding data from other years, but rather hoping to focus on cases in 2023 while supplementing cases from other years that are closely related to the topic, so as to fully understand the trends and background. The retrieval system can prioritize cases in 2023 while supplementing data from related years.
[0011] (2) Data sparsity within a specific time range is also an important issue.
[0012] For example, if a region has fewer public environmental pollution penalty cases in 2023, searching strictly by "2023" may result in insufficient results. In this case, relaxing the time limit (such as expanding to "2022" and "2023") can increase the coverage of search results and provide users with more valuable reference information. Summary of the invention
[0013] The present invention proposes a query expansion method based on generative retrieval, which aims to solve the problem of "zero results" or low-relevance search results caused by too many restrictive conditions in complex queries. The method uses a two-stage expansion process. In the first stage, a large language model is used to deeply analyze the query and extract keywords that can accurately express the user's intention; in the second stage, these keywords are used to guide the generation process, appropriately delete or weaken overly strict query conditions, and replace precise terms with broader terms, thereby improving the flexibility and coverage of the query. This method effectively improves the relevance and recall rate of query expansion, so that complex queries can obtain more accurate and comprehensive search results.
[0014] To achieve the above object, the present invention provides a query expansion method based on generative retrieval, comprising:
[0015] Extract keywords from the original query through a large language model to obtain core keywords;
[0016] After the core keywords are expanded, the keywords are replaced according to the context adjustment key words to generate an initial query expression;
[0017] Performing context analysis and constraint identification on the original query to identify and obtain strict query conditions;
[0018] After weakening and relaxing the strict query conditions, a target query expression is finally generated that meets the original query intent and covers more relevant results.
[0019] Preferably, the process of extracting keywords from the original query using a large language model to obtain core keywords includes:
[0020] The original query is deeply semantically understood through a large language model, and the model is guided by a prompt template to identify the most representative and high semantic density keywords in the original query to obtain the core keywords.
[0021] Preferably, the formula for obtaining the core keywords is:
[0022] K = Extract (Q, LLM, Prompt)
[0023] Among them: Q represents the original query; LLM represents the large language model; Prompt represents the prompt template, which is used to guide the model to understand the query and extract keywords.
[0024] Preferably, the process of expanding the core keywords includes:
[0025] Based on the semantic model or knowledge graph, the core keyword set K is expanded to generate an expanded keyword set K′.
[0026] Preferably, the formula for generating the expanded keyword set K′ is:
[0027] K′=K∪Expand(K)
[0028] Among them, Expand(K) represents an expansion function that can add synonyms, near-synonyms or field-related terms.
[0029] Preferably, the process of replacing keywords according to context adjustment includes:
[0030] For the expanded keyword set K′, a replacement operation is performed according to the context C to generate a replaced keyword set K″.
[0031] Preferably, the formula for generating the replaced keyword set K″ is:
[0032] K′ ′=Replace(K′ ,C) (3)
[0033] Replace(K′, C) means replacing some keywords with expressions that are more semantically consistent according to context information.
[0034] Preferably, the process of performing context analysis and constraint identification on the original query to identify and obtain strict query conditions includes:
[0035] Given the original query Q, we first analyze the query context C to identify whether there are overly strict query conditions;
[0036] Define the condition set in the query as Cstrict, and the context analysis process is expressed as:
[0037] Cstrict={c 1 , c 2 , c 3}
[0038] Among them, c 1 Indicates an overly specific time frame; c 2 Indicates an overly strict geographic location or physical characteristic; 3 Indicates that an exact keyword match exists.
[0039] Preferably, the process of weakening and relaxing the strict query condition includes:
[0040] For the identified condition set Cstrict, apply the condition identification algorithm to determine the conditions that need to be relaxed;
[0041] For each condition ci∈Cstrict, the decision function F is defined as:
[0042]
[0043] Then the strict condition set Cstrict is decomposed into two parts, and the condition set that needs to be relaxed is:
[0044]
[0045] The set of conditions that do not need to be relaxed is:
[0046]
[0047] For the query condition Cto_relax that needs to be relaxed, it is weakened. The formula expression is:
[0048] C′=Relax(Cto_relax) (8)
[0049] Among them: C′ represents the adjusted query condition; Relax (Cstrict) represents the process of weakening and relaxing the strict query condition, including: relaxing the time range, relaxing the geographical conditions, and relaxing the keywords.
[0050] Preferably, the process of generating a target query expression that satisfies the original query intent and covers more relevant results includes:
[0051] Generate the target query expression Q′ using the keyword set K″ and the adjusted constraint condition C′;
[0052] The formula expression is:
[0053] Q′=Generate(K″,C′)
[0054] Here, Generate(K″, C′) indicates generating a new query according to the keyword set K″ and the relaxed constraint condition C′.
[0055] Compared with the prior art, the present invention has the following advantages and technical effects:
[0056] The present invention provides a query expansion method based on generative retrieval, which can perform deep semantic understanding of queries and extract representative keywords through a large language model; optimize keyword sets using expansion and replacement strategies to generate a wider range of query expressions; and identify and weaken overly strict query conditions by analyzing the query context to ensure the flexibility and adaptability of the query. Through these technical features, the present invention can effectively expand the coverage of queries, improve the recall rate of queries, and avoid the "zero result" problem caused by too many restrictive conditions. At the same time, the method's condition weakening and keyword expansion strategy ensure that the query can cover more relevant information while still maintaining an accurate expression of the user's original intention. Finally, the present invention significantly improves the relevance and comprehensiveness of retrieval results and enhances the effectiveness and adaptability of queries by optimizing the query expansion process. BRIEF DESCRIPTION OF THE DRAWINGS
[0057] The drawings constituting a part of the present application are used to provide a further understanding of the present application. The illustrative embodiments and descriptions of the present application are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:
[0058] Figure 1 The figure is a schematic diagram of a method flow of an embodiment of the present invention. DETAILED DESCRIPTION
[0059] It should be noted that, in the absence of conflict, the embodiments and features in the embodiments of the present application can be combined with each other. The present application will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.
[0060] It should be noted that the steps shown in the flowcharts of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and that, although a logical order is shown in the flowcharts, in some cases, the steps shown or described can be executed in an order different from that shown here.
[0061] like Figure 1 As shown, this embodiment provides a query expansion method based on generative retrieval, including the following steps:
[0062] Extract keywords from the original query through a large language model to obtain core keywords;
[0063] After expanding the core keywords, the keywords are replaced according to the context adjustment to generate the initial query expression;
[0064] Perform context analysis and constraint identification on the original query to obtain strict query conditions;
[0065] After weakening and relaxing the strict query conditions, a target query expression is finally generated that meets the original query intent and covers more relevant results.
[0066] Further, step 1: keyword extraction;
[0067] The “extraction” operation in this process involves using a large language model to perform deep semantic understanding on the original query Q, and guiding the model to identify the most representative and high semantic density keywords in the original query through the prompt template. This process is expressed by formula (1):
[0068] K=Extract(Q, LLM, Prompt) (1)
[0069] Among them: Q represents the original query; LLM represents a large language model, such as the GPT series model; Prompt represents a carefully designed prompt template used to guide the model to better understand the query and perform keyword extraction.
[0070] Further, step 2: keyword expansion;
[0071] Based on the semantic model or knowledge graph, the core keyword set K is expanded to generate the expanded keyword set K′. This process is expressed by formula (2):
[0072] K′=K∪Expand(K) (2)
[0073] Among them, Expand(K) represents the expansion function, which increases the flexibility of the query by adding synonyms, near-synonyms or field-related terms.
[0074] Further, step 3: keyword replacement;
[0075] For the expanded keyword set K′, a replacement operation is performed according to the context C to generate a replaced keyword set K″. This process is expressed by formula (3):
[0076] K′ ′=Replace(K′ ,C) (3)
[0077] Replace(K′, C) means replacing some keywords with expressions that are more semantically consistent with the context information, such as replacing “mobile phone” with “mobile device”.
[0078] Further, step 4: context analysis and constraint identification;
[0079] Given the original query Q, we first analyze the query context C to identify whether there are overly strict query conditions. Assuming the condition set in the query is Cstrict, the context analysis process can be expressed as:
[0080] Cstrict={c 1 , c 2 , c 3} (4)
[0081] Where: c 1 Indicates an overly specific time range (e.g., "January 1, 2024 to December 31, 2024"); c 2 Indicates overly restrictive geographic locations or physical characteristics (such as "only products in Shanghai"); 3 Indicates that strict keyword matching exists (such as "match only 'Apple phone' and not 'mobile phone'").
[0082] Further, step 5: weakening and relaxing conditions
[0083] For the identified condition set Cstrict, a condition identification algorithm, such as rule-based pattern matching or machine learning methods, is applied to determine which conditions need to be relaxed. For each condition ci∈Cstrict, the following decision function F is defined:
[0084]
[0085] Then the strict condition set Cstrict can be decomposed into two parts:
[0086] Among them, the set of conditions that need to be relaxed is:
[0087]
[0088] The set of conditions that do not need to be relaxed is:
[0089]
[0090] For the query condition Cto_relax that needs to be relaxed, the following method is used to weaken it, which can be expressed as formula (8):
[0091] C′=Relax(Cto_relax) (8)
[0092] Where: C′ represents the adjusted query condition;
[0093] Relax (Cstrict) indicates the process of weakening and relaxing strict query conditions, including: relaxing time range, relaxing geographical conditions, and relaxing keywords.
[0094] Furthermore, the time range can be relaxed: if the query contains a very strict time interval, it may be weakened to "the last year" or "the past few months";
[0095] Relaxation of geographical conditions: If the query requires a precise geographical location, it may be expanded to a larger area or relaxed on a city basis;
[0096] Keyword relaxation: Add or replace some overly specific keywords to find broader matches. For example, expand "iPhone14" to "latest smartphones."
[0097] Further, step 6: generating a new query expression;
[0098] Using the keyword set K″ and the adjusted constraint condition C′, a more flexible and universal query expression Q′ is generated, which can be expressed as formula ():
[0099] Q′=Generate(K′ ′ , C′ ) (9)
[0100] Here, Generate(K″, C′) indicates generating a new query according to the keyword set K″ and the relaxed constraint condition C′.
[0101] The verification results of this embodiment on the data set DL20 are as follows.
[0102] The full name of the DL20 dataset is Deep Learning for Information Retrieval 2020. It is a benchmark dataset for information retrieval tasks, used to test and evaluate the performance of deep learning-based retrieval models on multiple query and document collections. The dataset contains real-world queries and related documents, and is used to promote the application and research of deep learning technology in the field of information retrieval.
[0103] The evaluation indicators on the DL20 dataset include:
[0104] Map@100: average precision of the first 100 documents;
[0105] Ndcg@10: Normalized discounted cumulative gain, ranking effect of the top 10 documents;
[0106] r@10: recall of the first 10 documents;
[0107] As shown in Table 1, the experimental results on the DL20 dataset show that the proposed method performs significantly better than other methods, especially achieving the highest values in the three evaluation indicators Map@100 (0.3068), Ndcg@10 (0.5130) and r@10 (0.2680).
[0108] Table 1
[0109]
[0110]
[0111] Here is a description of each model:
[0112] BM25: Calculates the relevance between query and document based on term frequency (TF), inverse document frequency (IDF) and document length. It is a classic document scoring function.
[0113] BM25+RM3: Combine BM25 and pseudo-relevance feedback (RM3) for query expansion, and rewrite the query by selecting keywords from highly relevant documents to improve retrieval accuracy.
[0114] GRF-Queries: Generate query variants using generative models and rewrite the original query to improve retrieval relevance.
[0115] GRF-Keywords: Extract keywords through generative models to construct new queries, focusing on core content.
[0116] The two-stage method proposed in this embodiment optimizes keyword extraction and performs query relaxation through a large language model, thereby improving the accuracy and recall rate of retrieval.
[0117] The verification results of this embodiment on the MSMARCO dataset are as follows:
[0118] The MSMARCO dataset is a large-scale collection designed to promote the application of deep learning techniques in information retrieval. The MSMARCO paragraph retrieval dataset contains more than 8.8 million paragraphs and more than 500,000 query-paragraph pairs for training. In addition, 6,980 queries are reserved for evaluation, which constitute the MSMARCO development set (Dev set).
[0119] The evaluation indicators on the MSMARCO dataset include:
[0120] Map@100: average precision of the first 100 documents;
[0121] r@10: recall of the first 10 documents;
[0122] As shown in Table 2, the experimental results on the MSMARCO dataset show that the proposed method performs significantly better than other methods, especially achieving the highest values in the MRR@10 (0.218) and r@10 (0.47) evaluation indicators.
[0123] Table 2
[0124] Model MRR@10 r@10 BM25 0.1875 0.3916 BM25+RM3 0.1646 0.374 BM25+Rocchio 0.1684 0.3769 Ours 0.2180 0.4700
[0125] Here is a description of each model:
[0126] BM25: Calculates the relevance between query and document based on term frequency (TF), inverse document frequency (IDF) and document length. It is a classic document scoring function.
[0127] BM25+RM3: Combine BM25 and pseudo-relevance feedback (RM3) for query expansion, and rewrite the query by selecting keywords from highly relevant documents to improve retrieval accuracy.
[0128] BM25+Rocchio: Combine BM25 with the Rocchio model, use BM25 to calculate the preliminary relevance score between the document and the query, and then use the Rocchio model to adjust the query based on the feedback to further improve the accuracy of the retrieval results
[0129] The two-stage method proposed in this embodiment optimizes keyword extraction and performs query relaxation through a large language model, thereby improving the accuracy and recall rate of retrieval.
[0130] As a supplementary embodiment, the specific implementation mode of the present invention is demonstrated below by taking "What are the symptoms of influenza in children in winter?" as an example.
[0131] Original query: Q = "How to prevent spring pollen allergies?"
[0132] Step 1: Keyword extraction;
[0133] In this step, the large language model (LLM) and the designed prompt template (Prompt) are used to semantically understand the original query Q and extract the keywords in the query to obtain the set K:
[0134] K = {Spring, pollen, allergy, prevention}
[0135] Step 2: Keyword expansion;
[0136] Next, the keyword set K is expanded based on the semantic model or knowledge graph, and synonyms, near synonyms, and related terms are used to enhance the flexibility of the query, generating the expanded keyword set K′:
[0137] K′={spring, pollen, allergy, prevention, hay fever, seasonal allergies, allergy protection, anti-allergy}
[0138] Step 3: Keyword replacement;
[0139] Based on the context C (e.g., a specific scene or context), the expanded keyword set K′ is replaced to generate a new keyword set K″. If there is a condition such as "spring" in the query, it can be replaced with "allergy high season" or "allergy season" according to the context to make the query more general or flexible. The replaced keyword set K″ is:
[0140] K″={allergy season, pollen, allergies, prevention, hay fever, seasonal allergies, allergy protection, anti-allergy}
[0141] Step 4: Context analysis and constraint identification;
[0142] Analyze the query context C and identify overly strict query conditions. Suppose we have the following strict condition set Cstrict:
[0143] Cstrict = {time range: spring, too strict keyword: pollen, too narrow geographical location: only in certain areas}
[0144] These conditions include:
[0145] Time frame: "Spring" is a narrower time frame.
[0146] The keyword "pollen" is too specific and may limit the scope of the search.
[0147] Geographic location: If the query conditions require focusing only on a specific region (for example, "Prevention measures for spring pollen allergies in Shanghai"), it is also a strict condition.
[0148] Step 5: Weakening and relaxing conditions;
[0149] According to the results of condition identification, weakening and relaxation rules are applied. The specific weakening process is as follows:
[0150] Relax the time range: Expand “spring” to “high allergy season” or “all year round”.
[0151] Keyword expansion: Expand "pollen" to "common allergens" or "allergens".
[0152] Relaxation of geographical location: expand "Shanghai" to "areas with high incidence of allergies" or "nationwide".
[0153] The relaxed condition set C′ is:
[0154] C′={allergy high season, common allergens, nationwide}
[0155] Step 6: Generate a new query expression;
[0156] Finally, based on the keyword set K″ and the relaxed condition set C′, a new query expression Q′ is generated:
[0157] Q′=“How to prevent allergic symptoms caused by common allergens during the allergy season?”
[0158] Through this method, the original query "How to prevent spring pollen allergies?" is transformed into a more flexible and universal query expression. These steps enable the query to cover a wider search scope while retaining the core intent of the original query.
[0159] This embodiment provides a method based on keyword extraction and context analysis, which extracts a set of keywords with high semantic density from the original query, and uses these keywords to guide the process of generating more flexible and universal query expressions; the strict conditions and context constraints in the query automatically weaken and relax query conditions such as time range and population restrictions to generate a query expression that meets the original query intent and covers more relevant results.
[0160] The above are only preferred specific implementations of the present application, but the protection scope of the present application is not limited thereto. Any changes or substitutions that can be easily thought of by a person skilled in the art within the technical scope disclosed in the present application should be included in the protection scope of the present application. Therefore, the protection scope of the present application should be based on the protection scope of the claims.
Claims
1. A query expansion method based on generative retrieval, characterized in that: include: Extract keywords from the original query through a large language model to obtain core keywords; After the core keywords are expanded, the keywords are replaced according to the context adjustment key words to generate an initial query expression; Performing context analysis and constraint identification on the original query to identify and obtain strict query conditions; After weakening and relaxing the strict query conditions, a target query expression is finally generated that meets the original query intent and covers more relevant results.
2. The method according to claim 1, characterized in that The process of extracting keywords from the original query using a large language model and obtaining core keywords includes: The original query is deeply semantically understood through a large language model, and the model is guided by a prompt template to identify the most representative and high semantic density keywords in the original query to obtain the core keywords.
3. The method according to claim 1, characterized in that The formula for obtaining the core keywords is: K = Extract (Q, LLM, Prompt) Among them: Q represents the original query; LLM represents the large language model; Prompt represents the prompt template, which is used to guide the model to understand the query and extract keywords.
4. The method according to claim 1, characterized in that The process of expanding the core keywords includes: Based on the semantic model or knowledge graph, the core keyword set K is expanded to generate an expanded keyword set K′.
5. The method according to claim 4, characterized in that The formula for generating the expanded keyword set K′ is: K′=K∪Expand(K) Among them, Expand(K) represents an expansion function that can add synonyms, near-synonyms or field-related terms.
6. The method according to claim 1, characterized in that The process of replacing keywords based on contextual adjustments includes: For the expanded keyword set K′, a replacement operation is performed according to the context C to generate a replaced keyword set K″.
7. The method according to claim 6, characterized in that The formula for generating the replaced keyword set K″ is: K′ ′=Replace(K′ ,C) (3) Replace(K′, C) means replacing some keywords with expressions that are more semantically consistent according to context information.
8. The method according to claim 1, characterized in that The process of performing context analysis and constraint identification on the original query to obtain strict query conditions includes: Given the original query Q, we first analyze the query context C to identify whether there are overly strict query conditions; Define the condition set in the query as Cstrict, and the context analysis process is expressed as: Cstrict = {c1, c2, c3} Among them, c1 represents an overly specific time range; c2 represents an overly strict geographic location or physical feature; and c3 represents the existence of a strict keyword match.
9. The method according to claim 1, characterized in that: The process of weakening and relaxing the strict query conditions includes: For the identified condition set Cstrict, apply the condition identification algorithm to determine the conditions that need to be relaxed; For each condition ci∈Cstrict, the decision function F is defined as: Then the strict condition set Cstrict is decomposed into two parts, and the condition set that needs to be relaxed is: The set of conditions that do not need to be relaxed is: For the query condition Cto_relax that needs to be relaxed, it is weakened. The formula expression is: C′=Relax(Cto_relax) (8) Among them: C′ represents the adjusted query condition; Relax (Cstrict) represents the process of weakening and relaxing the strict query condition, including: relaxing the time range, relaxing the geographical conditions, and relaxing the keywords.
10. The method according to claim 1, characterized in that The process of generating a target query expression that satisfies the original query intent and covers more relevant results includes: Generate the target query expression Q′ using the keyword set K″ and the adjusted constraint condition C′; The formula expression is: Q′=Generate(K″,C′) Here, Generate(K″, C′) indicates generating a new query according to the keyword set K″ and the relaxed constraint condition C′.
Citation Information
Cited By
Retrieval method, electronic equipment, medium and product
CN122470729A