Tobacco business information query rewriting method and system based on direct preference optimization
By constructing a tobacco business fuzzy question preference dataset and using a direct preference optimization algorithm to train a query rewriting model, the problems of user query ambiguity and standardization in tobacco government affairs scenarios are solved, efficient query rewriting and information retrieval are achieved, and the accuracy and standardization of tobacco business information queries are improved.
Patent Information
- Application Number
- CN202510790536.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-13
- Publication Date
- 2025-10-28
AI Technical Summary
In existing tobacco government affairs scenarios, user queries are vague and semantically unclear. Traditional query rewriting methods are difficult to adapt to the professional and standardized requirements of tobacco business, resulting in low retrieval recall rates. Existing technologies lack an effective model training mechanism to combine human expression preferences for query rewriting.
We construct a fuzzy question preference dataset for the tobacco business, train a query rewriting task model using the direct preference optimization algorithm, generate a query rewriting optimization model that conforms to the specifications, perform query rewriting using a large language model, and combine a domain terminology reinforcement module and the DPO preference optimization algorithm to directly learn human expression preferences and generate structured query expressions.
It has significantly improved the accuracy and standardization of tobacco business information queries, improved user experience, met the standardization requirements of government services, and improved the recall rate and accuracy of information retrieval.
Smart Images

Figure CN120849437A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of data processing technology, and in particular to a method and system for rewriting tobacco business information queries based on direct preference optimization. Background Technology
[0002] In terms of understanding and rewriting fuzzy queries, early systems mainly relied on rule templates and keyword matching strategies, lacking semantic understanding and unable to adapt to natural language input with high degrees of freedom. Subsequently, with the development of deep learning, researchers proposed query rewriting methods based on neural networks, such as the LSTM-based Seq2Seq architecture and the BERT2BERT or T5 model using the Transformer structure. These models map the original natural language query into a clearer and more structured expression through supervised learning, achieving good results in open-domain question answering and search engines. However, they still have three significant limitations in government affairs scenarios: (1) they rely heavily on large-scale manually labeled data, while government affairs corpora are scarce and labeling costs are high; (2) the training objective is only to optimize language similarity or grammatical integrity, ignoring the final retrieval effect; (3) the output expression often becomes templated, deviating from the preservation of the user's original intent and expression preferences.
[0003] To enhance the ability of queries to guide downstream retrieval, query representation expansion methods combining the capabilities of large language models have emerged in recent years. For example, the HyDE (Hypothetical Document Embedding) method uses a large language model to generate hypothetical segments for the query, then transforms them into vectors for semantic retrieval; the Query2Doc method generates pseudo-documents to construct a more complete query representation. While these methods have some effectiveness in open domains, in highly specialized, terminologically standardized tobacco information query scenarios, the assumptions are prone to generating knowledge illusions or semantic drift, affecting retrieval stability and accuracy. Meanwhile, addressing the issue of traditional query rewriting methods' insufficient perception of the final task's effect, some recent studies propose training rewriting models from the perspective of overall system performance. Among them, Salesforce AI proposed a framework for optimizing the performance of RAG (Retrieval Enhancement Generation) systems through query rewriting. It utilizes a large language model and a sorter to construct pseudo-label data, trains the rewriter, thereby improving the quality of retrieval recall and the final answer effect after rewriting. Furthermore, the INFO-RAG method proposed by the institution further defines the role of the large language model in the RAG system as an "information refiner" and enhances its ability to process incomplete or erroneous retrieved information through unsupervised training methods. These methods all demonstrate that improving the "quality of input expression" and the "model's ability to utilize retrieved information" in question-answering systems are key paths to improving the accuracy of the final response. However, the aforementioned methods mostly focus on the English open-domain environment and are primarily geared towards general-purpose language models or large-scale knowledge question-answering systems. They still lack adaptability to problems such as ambiguous Chinese expressions, non-standard question formulations, and scarce data in tobacco-related government affairs scenarios. Especially in the area of truly rewriting fuzzy query expressions, an effective mechanism that can incorporate human expression preferences for model training has not yet been established. User queries often have a gap between policy perception biases (such as the colloquial expression "Is it easy to get a tobacco license?") and the standardization of business terminology (requiring precise matching of "Application conditions for a tobacco retail license"), necessitating semantic alignment through query rewriting. When general rewriting techniques are applied directly, three major unique problems arise: first, the standardization of terminology is extremely demanding (e.g., "selling cigarettes" must be mapped to the name of a legal document); second, policy provisions vary geographically and in terms of timeliness (e.g., different provinces have different regulations regarding the spacing between retail outlets); and third, sensitive policy expressions must be avoided (e.g., correcting "find someone to handle it" to "legal application process"). Ordinary text rewriting techniques struggle to meet the requirements for accuracy and compliance. Summary of the Invention
[0004] The main objective of this application is to provide a method and system for rewriting tobacco business information queries based on direct preference optimization.
[0005] The technical solution adopted in this invention is:
[0006] On one hand, embodiments of the present invention provide a method for rewriting tobacco business information queries based on direct preference optimization, the method comprising the following steps:
[0007] Construct a dataset of preferences for fuzzy questions in the tobacco business;
[0008] Based on the aforementioned tobacco business fuzzy question preference dataset, a query rewriting task model is constructed;
[0009] The query rewriting task model is trained using a direct preference optimization algorithm to obtain a query rewriting optimization model;
[0010] Based on the query rewriting optimization model, the query rewriting of tobacco business information is completed.
[0011] Furthermore, the construction of the tobacco business fuzzy question preference dataset includes the following steps:
[0012] Construct the original fuzzy query list;
[0013] Construct a high-quality list of subqueries;
[0014] Construct a list of low-quality subqueries;
[0015] Based on the original fuzzy query list, the high-quality subquery list, and the low-quality subquery list, a database triplet is obtained;
[0016] Based on the database triples, a fuzzy question preference dataset for tobacco business is obtained.
[0017] Furthermore, constructing the original fuzzy query list includes the following steps:
[0018] We obtain user-asked tobacco-related questions from public consultation platforms and filter out vague expressions to obtain user fuzzy query data.
[0019] Simulated fuzzy questions are generated based on tobacco process specification information to obtain simulated question data;
[0020] By combining a large language model with process document segmentation, fuzzy questions are automatically generated to obtain model-generated query data.
[0021] Based on the user's fuzzy query data, the simulated question data, and the model-generated query data, an original fuzzy query list is obtained.
[0022] Furthermore, constructing the high-quality subquery list includes the following steps:
[0023] Based on the original fuzzy query list, the information of the original fuzzy query is transformed into matching statement information that conforms to preset standard conditions by rewriting.
[0024] Based on the original fuzzy query list, the large language model is guided by the instruction template to output structured subquery data;
[0025] Based on the matching statement information and the structured subquery data, a list of high-quality subqueries is obtained.
[0026] Furthermore, constructing the list of low-quality subqueries includes the following steps:
[0027] Based on the original fuzzy query list and the high-quality subquery list, statement information deviation data is generated using a large language model;
[0028] Based on the deviation of the data from the stated information, a list of low-quality subqueries is obtained.
[0029] Furthermore, the formula used to construct the query rewriting task model based on the tobacco business fuzzy question preference dataset includes:
[0030]
[0031] Where, q * The rewriting question generated by the query rewriting task model; Q is the fuzzy question preference dataset for the tobacco business; P θ Let θ be the conditional probability distribution of the candidate questions in the query rewriting task model under parameter θ; q0 is a fuzzy question.
[0032] Furthermore, the step of training the query rewriting task model using the direct preference optimization algorithm to obtain the query rewriting optimization model includes the following steps:
[0033] Initialize the model structure and parameters;
[0034] Based on the query rewrite task model and the model structure and parameters, the tobacco business fuzzy question preference dataset is loaded to obtain the reference model and the current strategy model.
[0035] Based on the reference model and the current policy model, the generation probability and the construction loss function are calculated using the direct preference optimization algorithm.
[0036] Based on the probabilities and the constructed loss function, obtain the updated policy model;
[0037] Based on the update strategy model, a query rewrite optimization model is obtained.
[0038] Furthermore, the formulas used to calculate the generation probability and construct the loss function based on the direct preference optimization algorithm according to the reference model and the current policy model include:
[0039]
[0040] Among them, L DPO (π θ ;π ref ) represents the loss function; π θ For the current strategy model; π ref Here is the reference model; x is the user's original fuzzy query; y1 is the first rewritten candidate sentence; y2 is the second rewritten candidate sentence; y w Rewritten for higher quality; y l This is a poor rewrite; σ is the Sigmoid function; β is the temperature hyperparameter; π * For the optimal strategy model, E is the expected value; D is the data distribution.
[0041] On the other hand, embodiments of the present invention also provide a tobacco business information query rewriting system based on direct preference optimization, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the tobacco business information query rewriting method based on direct preference optimization as described above.
[0042] On the other hand, embodiments of the present invention also provide a computer-readable storage medium storing computer-executable instructions for causing a computer to execute the tobacco business information query rewriting method based on direct preference optimization as described above.
[0043] The embodiments of this application include at least the following beneficial effects: This application provides a method and system for rewriting tobacco business information queries based on direct preference optimization. The steps of this invention include constructing a tobacco business fuzzy question preference dataset; constructing a query rewriting task model based on the tobacco business fuzzy question preference dataset; training the query rewriting task model using a direct preference optimization algorithm to obtain a query rewriting optimization model; and completing the rewriting of tobacco business information queries based on the query rewriting optimization model. This invention can effectively compensate for the poor adaptability of traditional question-answering systems to non-standard expressions, significantly improve the system's understanding and response accuracy to natural language input, enhance the processing effect of tobacco business information, and accurately generate standardized tobacco business information for user fuzzy queries. Attached Figure Description
[0044] Figure 1 This is a schematic diagram of the tobacco business information query rewriting method based on direct preference optimization provided in an embodiment of the present invention;
[0045] Figure 2 This is a schematic diagram illustrating the process of constructing a fuzzy question preference dataset for tobacco business information provided in an embodiment of the present invention;
[0046] Figure 3 This is a schematic diagram of the prompt word template and specific process provided in the embodiments of the present invention;
[0047] Figure 4 This is a schematic diagram of a specific prompt word template provided in an embodiment of the present invention;
[0048] Figure 5 This is a flowchart of the low-quality subquery construction process provided in an embodiment of the present invention;
[0049] Figure 6 This is a schematic diagram of the training process of the query rewrite optimization model provided in the embodiment of the present invention. Detailed Implementation
[0050] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative of this application and are not intended to limit it. In the following description, when referring to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with those of this application; they are merely examples of apparatuses and methods consistent with some aspects of the embodiments of this application as detailed in the appended claims.
[0051] It is understood that the terms “first,” “second,” etc., used in this application may be used herein to describe various concepts, but unless otherwise stated, these concepts are not limited by these terms. These terms are only used to distinguish one concept from another. For example, without departing from the scope of the embodiments of this application, first information may also be referred to as second information, and similarly, second information may also be referred to as first information. Depending on the context, the words “if,” “when,” or “in response to a determination” as used herein may be interpreted as “when…” or “when…” or “in response to a determination.”
[0052] As used in this application, the terms "at least one", "multiple", "each", "any", etc., "at least one" includes one, two or more, "multiple" includes two or more, "each" refers to each of the corresponding multiples, and "any" refers to any one of the multiples.
[0053] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which this application pertains. The terms used herein are for the purpose of describing the embodiments of this application only and are not intended to limit this application.
[0054] Before providing a detailed description of the embodiments of this application, some of the nouns and terms involved in the embodiments of this application will be explained first. The nouns and terms involved in the embodiments of this application are subject to the following interpretations.
[0055] 1) LSTM (Long Short-Term Memory), a type of recurrent neural network;
[0056] 2) Seq2Seq (Sequence-to-Sequence), a sequence-to-sequence model;
[0057] 3) Transformer, a model based on the self-attention mechanism;
[0058] 4) BERT2BERT model, a bidirectional Transformer to bidirectional Transformer model, based on the encoder-decoder structure of BERT, used for generation tasks;
[0059] 5) T5 model (Text-to-Text Transfer Transformer);
[0060] 6) HyDE (Hypothetical Document Embedding), a technique that enhances search performance by generating hypothetical documents;
[0061] 7) Query2Doc: Generates the document retrieved;
[0062] 8) Salesforce AI, the Salesforce artificial intelligence platform;
[0063] 9) RAG (Retrieval-Augmented Generation), a retrieval-enhanced generation method;
[0064] 10) INFO-RAG (Information Enhanced RAG), an improved version of RAG;
[0065] 11) The Sigmoid function is an activation function that maps the input to the (0,1) interval and is used for binary classification or probability output.
[0066] 12) DPO algorithm (Direct Preference Optimization);
[0067] 13) Chunk, text segmentation;
[0068] 14) LLaMA 3-7B, a large open-source model with 7 billion parameters developed by Meta;
[0069] 15) Qwen2.5-7B, the 7 billion parameter model of Ali Tongyi Qianwen 2.5 version;
[0070] 16) DeepSeek-V3-Instruct-7B, a fine-tuning model for instructions with 7 billion parameters developed by DeepSeek.
[0071] 17) RLHF (Reinforcement Learning from Human Feedback);
[0072] 18) SFT model (Supervised Fine-Tuning);
[0073] 19) Bradley-Terry, a probabilistic model for pairwise comparisons of a set of options;
[0074] 20) KL constraint (Kullback-Leibler Divergence Constraint): In reinforcement learning, the KL divergence constraint limits the difference between the policy model and the reference model to prevent excessive deviation.
[0075] 21) TODO: Compare loss functions;
[0076] 22) Instruct model, instruction fine-tuning model.
[0077] This invention addresses the common issues of ambiguity, colloquialism, and information gaps in user queries during tobacco business consultation scenarios, proposing a professional query rewriting method. In practical applications, ordinary users often use non-standard expressions such as "What are the requirements for applying for a tobacco license?" or "Is it easy to apply for a tobacco license?" These queries suffer from problems such as missing core terminology (failure to distinguish between retail and production licenses), ambiguous intent (failure to specify consultation conditions, procedures, or materials), and neglect of regional characteristics (failure to reflect local policy differences), directly leading to low retrieval rates in government Q&A systems. Through the domain terminology enhancement module and DPO preference optimization algorithm designed in this invention, fuzzy queries can be accurately rewritten into structured expressions such as "What legal conditions must be met to apply for a tobacco retail license (retail category)?" This effectively filters out suggestive expressions like "using connections," satisfying the standardization requirements of government services while significantly improving the user experience. This technology solves the unique challenges of tobacco business, such as the timeliness of policies, significant regional differences, and specialized terminology, providing reliable technical support for intelligent government services.
[0078] The embodiments of the present invention will be further described below with reference to the accompanying drawings.
[0079] On one hand, embodiments of the present invention provide a method for rewriting tobacco business information queries based on direct preference optimization, referring to... Figure 1 The method for rewriting tobacco business information queries based on direct preference optimization includes the following steps:
[0080] S100. Construct a dataset of preferences for fuzzy questions related to tobacco business.
[0081] S200. Based on the tobacco business fuzzy question preference dataset, construct a query rewriting task model;
[0082] S300. The query rewriting task model is trained using the direct preference optimization algorithm to obtain the query rewriting optimization model;
[0083] S400. Based on the query rewrite optimization model, complete the query rewrite of tobacco business information.
[0084] The present invention discloses S100 for constructing a tobacco business fuzzy question preference dataset, which includes the following steps:
[0085] S110. Construct the original fuzzy query list;
[0086] S120. Construct a high-quality list of subqueries;
[0087] S130, Construct a list of low-quality subqueries;
[0088] S140. Based on the original fuzzy query list, the high-quality subquery list, and the low-quality subquery list, obtain the database triples;
[0089] S150. Based on the database triples, obtain the tobacco business fuzzy question preference dataset.
[0090] The S110 method for constructing the original fuzzy query list disclosed in this embodiment of the invention includes the following steps:
[0091] S111. Obtain user-asked tobacco information from public consultation platforms and filter out fuzzy expressions to obtain user fuzzy query data;
[0092] S112. Generate simulated fuzzy questions based on tobacco process specification information to obtain simulated question data;
[0093] S113. By combining a large language model with process document segmentation, fuzzy questions are automatically generated to obtain model-generated query data.
[0094] S114. Generate query data based on user fuzzy query data, simulated question data, and model to obtain the original fuzzy query list.
[0095] The S120 method for constructing a high-quality subquery list disclosed in this embodiment of the invention includes the following steps:
[0096] S121. Based on the original fuzzy query list, the information of the original fuzzy query is transformed into matching statement information that conforms to the preset standard conditions by rewriting.
[0097] S122. Based on the original fuzzy query list, guide the large language model to output structured subquery data through the instruction template;
[0098] S123. Based on the matching statement information and structured subquery data, obtain a list of high-quality subqueries.
[0099] As an optional implementation, in step S123 of this embodiment of the invention, a mapping relationship between fuzzy queries and structured subqueries is established by extracting keywords (e.g., changing "tobacco license" to "tobacco retail license") and classifying intent (e.g., changing "time" to "approval time limit").
[0100] The S130 method for constructing a low-quality subquery list disclosed in this embodiment of the invention includes the following steps:
[0101] S131. Based on the original fuzzy query list and the high-quality subquery list, generate statement information deviation data through a large language model;
[0102] S132. Based on the deviation of the data from the statement information, obtain a list of low-quality subqueries.
[0103] The S200 method disclosed in this embodiment of the invention constructs a query rewriting task model based on a tobacco business fuzzy question preference dataset, and the formulas used include:
[0104]
[0105] Where, q * The rewrite question generated by the query rewrite task model; Q is the tobacco business fuzzy question preference dataset; P θ The conditional probability distribution of candidate questions for the query rewrite task model under parameter θ; q0 is a fuzzy question.
[0106] The S300 method disclosed in this embodiment of the invention uses a direct preference optimization algorithm to train a query rewriting task model to obtain a query rewriting optimization model, including the following steps:
[0107] S310. Initialize the model structure and parameters;
[0108] S320. Based on the query rewrite task model and model structure and parameters, load the tobacco business fuzzy question preference dataset to obtain the reference model and the current strategy model.
[0109] S330. Based on the reference model and the current policy model, calculate the generation probability and construct the loss function using the direct preference optimization algorithm;
[0110] S340. Based on the probability and construct the loss function, obtain the updated policy model;
[0111] S350. Based on the update strategy model, the query rewrite optimization model is obtained.
[0112] The embodiment of this invention discloses S330, which calculates the generation probability and construction loss function based on the direct preference optimization algorithm according to the reference model and the current policy model. The formulas used include:
[0113]
[0114] Among them, L DPO (π θ ;π ref ) represents the loss function; π θ For the current strategy model; π ref Here is the reference model; x is the user's original fuzzy query; y1 is the first rewritten candidate sentence; y2 is the second rewritten candidate sentence; y w Rewritten for higher quality; y l This is a poor rewrite; σ is the Sigmoid function; β is the temperature hyperparameter; π * For the optimal strategy model, E is the expected value; D is the data distribution.
[0115] As an optional implementation, the embodiments of the present invention aim to solve the problem of insufficient applicability of existing query rewriting methods in question-answering systems, specifically including problems such as weak understanding of fuzzy queries, rewriting results not conforming to user expression habits, and deviations in generated results. The present invention proposes a fine-tuning method for tobacco business information query rewriting model based on human preference optimization method, which improves the parsability, retrieval, and response accuracy of fuzzy queries in actual application information review business.
[0116] This method constructs a query rewriting model based on the Direct Preference Optimization (DPO) algorithm. By introducing human preference ranking data, the model can directly learn human expression preferences and rewriting criteria in actual business scenarios. Compared with existing technologies that rely on generating pseudo-documents or hypothetical documents, this invention uses real user original queries and human rewrite pairs as the basis for constructing training data. By annotating multiple candidate rewrite results with preferences, paired training samples are generated, which significantly improves the expressive authenticity of the training data and the target alignment of the training signal, avoiding semantic inconsistencies caused by model illusions or generation biases.
[0117] During the model optimization phase, this invention employs the DPO algorithm to directly optimize the model output. Compared to previous strategies that required training a reward model or indirectly modeling preferences, this method achieves direct learning of human preferences in a simpler and more efficient way, improving the expressive accuracy and structural regularity of query rewriting. This training method effectively reduces linguistic redundancy and structural ambiguity in the query rewriting process, improves the matching degree between query statements and documents in the knowledge base, and thus enhances the accuracy of retrieval recall.
[0118] This method eliminates the need for iterative retrieval or pseudo-document generation modules in its structural design, simplifying the overall system flow and reducing computational complexity and deployment costs. In practical applications, the rewritten model can be integrated as a pre-module into the tobacco information intelligent question-answering system to structurally rewrite ambiguous, abstract, or ambiguous questions posed by users, effectively improving the performance of subsequent information retrieval and answer generation modules.
[0119] This invention differs from existing query rewriting methods based on pseudo-document generation, prompting engineering, or supervised learning paradigms in its data construction method, training optimization strategy, and system integration approach. It boasts higher expression accuracy, stronger preference alignment capability, and superior system practicality, making it suitable for diverse and complex platform business query scenarios.
[0120] As an optional implementation, embodiments of the present invention include:
[0121] I. Method for Constructing a Fuzzy Question Preference Dataset for Tobacco Business
[0122] This invention provides a method for constructing a preference dataset for training a query rewriting task model, specifically for intelligent question-answering scenarios related to tobacco certificate application information. The method aims to construct a triplet structure data containing fuzzy queries (Prompt), high-quality subqueries (Chosen), and low-quality subqueries (Rejected), to drive a large language model to master high-quality query reconstruction strategies through contrastive supervised learning. The constructed dataset conforms to human preference orientation, significantly improving the system's ability to resolve fuzzy user questions and enhancing the accuracy and retrievability of query generation. The process for constructing the fuzzy question preference dataset for tobacco business information is as follows: Figure 2 As shown.
[0123] The construction of the original fuzzy query list includes three main sources. First, by searching for keywords such as "license" and "application" on authoritative public platforms like the State Tobacco Monopoly Administration's official website, publicly available public questions and answers are obtained. Questions with vague characteristics such as incomplete reference, semantic generalization, and colloquial tone are selected as the original fuzzy queries. These questions often do not explicitly specify the type of document, application process, or eligibility requirements, reflecting the natural expression of real users unfamiliar with policies. Second, based on policy and regulatory texts, vaguely expressed questions that might appear in actual consultations are drafted using human interpretation, such as "Is it difficult to prepare application information for a tobacco license?" or "Can application information for a tobacco retail license be generated online?" These questions are semantically ambiguous and incomplete, but are common in real-world public inquiry scenarios. Third, a large language model is used for assisted generation. Specifically, policy and regulatory documents are pre-processed in a structured manner, divided into several semantically complete content chunks. Each chunk is considered a potential "response content," and prompt words are constructed to guide the language model in generating corresponding fuzzy natural language questions. The prompts explicitly require the model to generate concise, general, conversational questions that avoid jargon. For example, "Generate a vague and concise question based on the following, avoiding jargon and emphasizing the user's perspective." This method can batch-construct natural-looking, realistically expressed vague questions, simulating the questioning habits of ordinary users with limited knowledge. The prompt template and specific process are as follows... Figure 3 As shown.
[0124] The construction of a high-quality subquery list focuses on transforming raw fuzzy queries into semantically clear, structurally standardized question expressions suitable for information retrieval. This process includes two paths: manual rewriting and model generation, ensuring that the generated results have a clear intent in their expression and cover the key information involved in the original question in terms of content. The manual rewriting path involves annotation personnel with an understanding of tobacco licensing policies manually writing subqueries according to established specifications. The goal of rewriting is to break down the user's fuzzy expression into a set of logically sound, thematically focused sub-questions that are easy to match with information. For example, the original fuzzy question "Is it difficult to obtain a tobacco license?" can be reconstructed into sub-questions such as "What are the application procedures for a tobacco monopoly license?" and "What conditions need to be met to apply for a tobacco monopoly license?" These sub-questions are semantically clear, highly targeted, and more suitable for driving the question-answering system to perform knowledge matching and answer generation.
[0125] In the model generation path, prompt words with instructions and constraints are constructed to guide the large language model to output well-structured and sufficiently comprehensive high-quality subqueries according to established standards. The prompt word template explicitly requires the model to generate several specific questions around the original fuzzy question, emphasizing that the content should be targeted, concise, and search-adaptable. For example, using prompt words such as "Transform the following fuzzy tobacco business-related questions into one or more clear sub-questions, ensuring that the questions are specific and clear, and revolve around the core intent" can effectively improve the quality and consistency of the generated questions. The generated results must be manually reviewed to filter out non-compliant questions, retaining semantically complete and logically consistent subqueries as the "chosen" part for constructing the preference dataset. Specific prompt word templates are as follows: Figure 4 As shown in the flowchart.
[0126] The construction of the low-quality subquery list focuses on generating semantically related but poorly expressed and retrieved questions. These questions serve as the "rejected" part in the construction of preference data, providing the model with comparative supervision signals and enhancing its ability to distinguish between high-quality and low-quality rewritten expressions. This process primarily relies on a large language model to generate content guided by prompt words. By constructing specific prompt word templates, the model is guided to output questions with a colloquial style, strong subjectivity, unclear logical structure, or information deviation, based on the original fuzzy question and the generated high-quality subqueries.
[0127] The prompt word template requires that the generated questions still maintain a certain semantic relevance to the original fuzzy query, but inaccurate or one-sided information, grammatically unnatural expressions, or even slight deviations from the topic can be intentionally introduced to form a comparative sample. For example, for the fuzzy question "Is it difficult to get a tobacco license?", and the corresponding high-quality subqueries "What are the application procedures for a tobacco monopoly license?" and "What conditions need to be met to apply for a tobacco monopoly license?", the generated low-quality subqueries might be "Has anyone's application information been stuck for several months?" or "I heard the process is extremely complicated, is that true?". Although these questions still revolve around the topic of "application business", they are defined as low-quality queries because they are highly subjective, lack specificity, and are unclear in expression, making it difficult for the system to complete an effective retrieval and answer.
[0128] The "rejected" data generated in the above manner not only possesses semantic relevance and formal diversity, but also provides discriminatory yet genuine negative samples during training. This helps the model more accurately grasp the expressive features that a "good query reconstruction" should possess during contrastive learning. The prompt word template is as follows: Figure 5 The flowchart for constructing low-quality subqueries is shown below.
[0129] To further enhance the model's generalization ability, this invention also introduces input augmentation mechanisms for fuzzy queries, mainly including back-translation and word order perturbation. The former generates semantically equivalent but formally different questions by translating the fuzzy question into English and then back-translating it into Chinese; the latter expands the diversity of fuzzy queries by adjusting word order or replacing words with synonyms. These augmented samples, while maintaining the original intent, help the model to be more robust to changes in input expression.
[0130] In terms of data organization, the final dataset is stored in JSON format. Each data sample contains three parts: prompt (the original list of fuzzy queries), chosen (a list of high-quality subqueries), and rejected (a list of low-quality subqueries). This triplet structure is suitable for preference optimization training paradigms and can be used for supervised fine-tuning of large language models using algorithms such as Direct Preference Optimization (DPO).
[0131] Through the above construction method, this invention provides a preference dataset generation scheme that is systematic, highly automated, and adaptable to actual business needs. It provides high-quality training corpus support for fuzzy query parsing tasks and can significantly improve the performance of intelligent question answering systems in tobacco license application business scenarios.
[0132] II. Principles and Implementation of Query Rewrite Optimization Model Training Method
[0133] In actual consultation scenarios for inquiring about application conditions for tobacco monopoly administrative licenses, user questions are often vague, semantically unclear, or lack key information, resulting in less than ideal results from keyword matching or semantic vector retrieval. To improve the system's ability to understand and respond to such vague questions, this invention proposes introducing a query rewriting mechanism. This involves rewriting the original fuzzy query using a model to generate a standard query sentence with greater semantic clarity and search relevance, facilitating downstream document matching and answer generation.
[0134] The query rewriting task can be formally defined as follows: Let q0 be the fuzzy question input by the user, whose semantic expression is incomplete or ambiguous. The goal of the query rewriting task model is to generate or select a rewriting question q from a given candidate space. * , making q * This approach can more accurately express the user's true intent and achieve better recall and matching results in actual knowledge base retrieval. Specifically, this task can be modeled as a condition generation problem:
[0135]
[0136] Where Q represents the set of all possible rewritten questions (tobacco business fuzzy question preference dataset), P θ To query the rewrite model, we need to determine the conditional probability distribution of candidate questions under parameter θ.
[0137] To optimize this objective, this invention further introduces a preference-based training method, which models the query rewriting task as a sequence generation process with subjective preference signals.
[0138] 1. Query rewrite model structure and input / output definition
[0139] The query rewriting module proposed in this invention is built upon current mainstream general-purpose large language model architectures, such as lightweight instruction fine-tuning models like LLaMA3-7B, Qwen2.5-7B, and DeepSeek-V3-Instruct-7B. These models possess excellent language understanding and instruction following capabilities, enabling them to transform fuzzy queries into structured queries with only natural language prompts. Compared to traditional information extraction or rule matching methods, large language models offer significant advantages in understanding intent ambiguity, completing omitted information, and generating standardized expressions, making them particularly suitable for the rewriting scenario in this invention, which deals with a large number of fuzzy business queries.
[0140] The model uses a standard instruction dialogue format as its input and output interface, uniformly organized into a multi-turn dialogue structure, containing three role fields: system, user, and assistant. The system field sets the model's behavioral instructions, such as "You are a query rewriting assistant, tasked with rewriting the user's fuzzy questions into clearer, more precise, and structured query expressions for subsequent information retrieval." The user field is used to input the user's natural language fuzzy query, such as "I want to ask what application materials are needed for an individual business owner to apply for a tobacco license?" The assistant field represents the standardized rewritten result generated by the model, such as "What materials are needed for an individual business owner to apply for a tobacco retail license?" This format is compatible with the calling interfaces of current mainstream open-source Instruct models and facilitates rapid integration during subsequent training and deployment.
[0141] In actual deployment, this module calls the large language model service in prompt mode and dynamically inserts user queries using preset prompt word templates to generate corresponding standard questions. To improve the consistency and stability of the model, all training data is organized with the same structure and uniformly formatted as the above three types of field combinations.
[0142] 2. Training methods for query rewriting optimization models
[0143] The training method for the query rewriting optimization model in this invention is based on human preference comparison data. It guides the model to learn and generate rewriting results that better match the target behavior through a reinforcement learning paradigm called Direct Preference Optimization (DPO). Compared to traditional reinforcement learning fine-tuning methods (RLHF), the DPO method eliminates explicit reward model learning and complex reinforcement learning training processes, greatly simplifying system implementation and reducing training resource overhead. It is suitable for fine-tuning tasks in the current large language model fine-tuning stage, and is particularly suitable for scenarios where subjective evaluation of rewriting results exists, as described in this invention. The training process of the query rewriting optimization model in this invention is as follows: Figure 6 As shown.
[0144] The DPO (Direct Preference Optimization) method takes human preference pairs as its core and defines the loss function directly in the policy space, thereby bypassing the reward model training and reinforcement learning process. It has advantages such as simplicity of implementation and stable training.
[0145] DPO no longer constructs an explicit reward function, but instead uses human preferences to apply the reward function to samples (x, y). + ,y - Direct optimization strategy for language models π θ This makes it more inclined to generate rewritten results y with higher preferences. + The theoretical basis lies in the fact that in preference probability modeling, there is a one-to-one mapping relationship between the optimal policy and the reward function, which can transform the optimization problem on the reward function into the minimization problem of the loss function on the policy probability.
[0146] Assume that an embodiment of the present invention has a preference dataset. For each original input x, y + For better rewriting results based on human preferences, y - This is a poor result. The embodiments of the present invention use π. θ This represents the policy model being trained, represented by π. ref This represents a fixed reference model (usually an SFT model). Under the Bradley-Terry assumptions, the preference distribution is:
[0147]
[0148] The key to DPO lies in expressing the reward r(y) as a function of the policy probability and the reference policy probability:
[0149] r(y) ∝ log(π) θ (y|x))-log(π ref (y|x))
[0150] The objective function of DPO is derived as follows: Starting from the same reinforcement learning objective as the traditional RLHF method, i.e., maximizing the expected value of the objective function:
[0151]
[0152] It can be proven that the optimal solution for the KL-constrained reward maximization objective in the equation takes the following form:
[0153]
[0154] in It is a separating function. Although a near-true reward function r can be obtained through maximum likelihood estimation. * Maximum likelihood estimate r Φ However, since Z(x) is still difficult to solve accurately, this expression presents certain difficulties in practical applications. Therefore, we can... Perform reparameterization so that it depends only on the optimal policy π. r Reference strategy π ref The segmentation function Z(.) of the position is used to indirectly represent the reward function.
[0155] right Taking the logarithm of both sides and performing algebraic derivation, we obtain:
[0156]
[0157] The above reparameterization method also applies to the real reward r. * and its corresponding optimal model π * Considering that the Bradley-Terry model relies only on the reward difference between two candidate outputs, i.e., p * (y1 f y2|x)=σ(r * (x,y1)-r * Substituting this reparameterized form (x, y2) into the preference modeling formula, the separating function Z(x) cancels out in the numerator and denominator, allowing the human preference probability to be expressed solely as the optimal policy π. * and reference strategy π ref The relationship between them.
[0158] Therefore, under the assumptions of the Bradley-Terry model, the optimal RLHF policy π * Satisfying Preference Model:
[0159]
[0160] Based on the above results, a strategy model π can be further constructed. θ The maximum likelihood training objective is used, thus bypassing the explicit reward modeling process. The final DPO training objective function is:
[0161]
[0162] The objective function implicitly fits the reward signal and directly optimizes the policy model parameters, causing it to converge to the optimal policy π. θ .
[0163] The training phase primarily relies on the preference question pair dataset constructed earlier as training samples. Each sample contains an original fuzzy user question and two rewrite candidates (labeled "better" and "poorer," respectively). These preference pairs are obtained through manual evaluation or rule-based selection, truly reflecting the preference for clear, standardized, and easily searchable question expressions in business needs. The DPO training objective is not simply to minimize the difference between the predicted rewrite and the "better" question, but rather to optimize a contrastive loss function (TODO) so that the model tends to generate expressions that are closer to human preference judgments during generation.
[0164] Instead of directly optimizing the rewritten text itself, DPO learns a policy model that, given the same user input, generates a higher probability of rewriting "better" questions compared to the base model. The training process relies on a fixed reference model (usually the Instruct model before fine-tuning) and a policy model being trained. By calculating the generation probabilities of "better" and "worse" questions for both models separately and constructing a contrastive loss function, the policy model is pushed to update in the direction of human preferences.
[0165] In the implementation of this invention, the query rewriting model is trained based on the DPO algorithm, and its detailed process includes the following key steps:
[0166] Step 1. Initialize Model Structure and Parameters: Load the base version of the large language model, which will serve as the initial policy model and reference model during training. Generally, the reference model can be directly initialized as the supervised fine-tuning (SFT) model π. SFT This serves as a source strategy for human preference data, ensuring that the strategy distribution is consistent with the preference data generation process and reducing training interference caused by distribution shift.
[0167] Step 2. Load the preference training dataset: Read the constructed preference question pair dataset. This dataset consists of a large number of triplet samples. Each sample contains a user's original fuzzy query x, and two rewritten candidate sentences y1 and y2, and the superior y1 has been labeled according to the method described above. w and the worse ones y l This results in a structure D = {(x,y} w ,y l The numbered pair format dataset.
[0168] Step 3. Calculate the generation probability and construct the loss function: For each sample, use the reference model π. ref Compared with the current policy model π θ For the two candidate sentences y w and y l Calculate its generation probability under input x, and further convert it into a log-likelihood probability value. Based on the DPO algorithm, optimize using the following comparative loss function:
[0169]
[0170] Among them, L DPO (π θ ;π ref ) represents the loss function; π θ For the current strategy model; π ref Here is the reference model; x is the user's original fuzzy query; y1 is the first rewritten candidate sentence; y2 is the second rewritten candidate sentence; y w Rewritten for higher quality; y l This is a poor rewrite; σ is the Sigmoid function; β is the temperature hyperparameter; π * For the optimal strategy model, E is the expected value; D is the data distribution.
[0171] The goal of optimizing the loss function is to enhance the model so that, when faced with the same input, it prioritizes generating outputs that are closer to human preferences.
[0172] Step 4. Policy Model Parameter Update: Based on the above loss function, update the policy model π using backpropagation. θ The parameters are updated. This process is performed in mini-batches until all training samples have been traversed, completing one training epoch.
[0173] Step 5. Reference Strategy Update: If the preference data does not originate from the currently initialized π ref It can be based on the optimal output sample (x,y) w ) for π ref Perform log-likelihood maximization fine-tuning to reduce π. ref The distribution difference between the actual preference data generation strategy and the distribution difference improves training stability.
[0174] Step 6. Validation and Early Stopping Mechanism: After every few rounds of training, evaluate the performance of the current policy model on the quality of generation and its matching degree with human preferences using an independent validation set. If the performance is found to be stabilizing or starting to decline, terminate training through the early stopping mechanism to avoid overfitting.
[0175] Compared with existing methods for keyword matching or static template rewriting in tobacco-related government affairs Q&A systems, this invention has the following technical advantages and beneficial effects:
[0176] 1. Enhanced ability to resolve fuzzy questions: This invention introduces a query rewriting mechanism based on a large language model for the first time, which can transform user-inputted fuzzy, generalized, and non-standard questions into structured and searchable standard queries. This effectively makes up for the poor adaptability of traditional question answering systems to non-standard expressions and significantly improves the system's ability to understand and respond to natural language input.
[0177] 2. Adapting to user preferences and enhancing generation quality: By constructing a dataset containing triples of "original query - high-quality rewrite - low-quality rewrite", the model training process uses human preferences as direct supervision signals, which significantly improves the accuracy, standardization, and user acceptance of generated queries, thereby optimizing the downstream knowledge retrieval and answer generation effects.
[0178] 3. Highly efficient and stable training, lower implementation threshold: Compared with the traditional reward-based RLHF method, this invention adopts the Direct Preference Optimization (DPO) strategy for training, which avoids the influence of reward modeling errors. It has the advantages of more stable training process, faster convergence speed and lower deployment cost, and is particularly suitable for rapid deployment in government low-resource environments.
[0179] 4. Supports offline training and flexible and convenient deployment: The preference data used can be uniformly constructed in the early stage, supporting offline training and iterative updates. No online interactive annotation is required, which greatly reduces the system construction and operation and maintenance costs, and facilitates rapid application in actual business scenarios such as tobacco administration systems.
[0180] 5. Strong model compatibility and good platform adaptability: The method of this invention can be adapted to mainstream open source large language models such as Qwen, LLaMA, and DeepSeek. It has good structural universality and platform independence, and can be flexibly deployed in different computing environments according to business needs, meeting the intelligent question answering needs of diverse government application scenarios.
[0181] On the other hand, embodiments of the present invention also provide a tobacco business information query rewriting system based on direct preference optimization, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the tobacco business information query rewriting method based on direct preference optimization as described above.
[0182] The processor and memory can be connected via a bus or other means. Memory, as a non-transitory computer-readable storage medium, can be used to store non-transitory software programs and non-transitory computer-executable programs. Furthermore, memory may include high-speed random access memory, and may also include non-transitory memory, such as at least one disk storage device, flash memory device, or other non-transitory solid-state storage device. In some embodiments, memory may optionally include memory remotely located relative to the processor, which can be connected to the processor via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof.
[0183] On the other hand, embodiments of the present invention also provide a computer-readable storage medium storing computer-executable instructions for causing a computer to execute the tobacco business information query rewriting method based on direct preference optimization as described above.
[0184] It will be understood by those skilled in the art that all or some of the steps and systems in the methods disclosed above can be implemented as software, firmware, hardware, and suitable combinations thereof. Some or all of the physical components can be implemented as software executed by a processor, such as a central processing unit, digital signal processor, or microprocessor, or as hardware, or as an integrated circuit, such as an application-specific integrated circuit. Such software can be distributed on a computer-readable medium, which can include computer storage media (or non-transitory media) and communication media (or transient media). As is known to those skilled in the art, the term computer storage media includes volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing information (such as computer-readable instructions, data structures, program modules, or other data). Computer storage media includes, but is not limited to, RAM, ROM, EEPROM, flash memory or other memory technologies, CD-ROM, digital versatile disc (DVD) or other optical disc storage, magnetic cartridges, magnetic tape, disk storage or other magnetic storage devices, or any other medium that can be used to store desired information and is accessible to a computer. Furthermore, as is well known to those skilled in the art, communication media typically embodies computer-readable instructions, data structures, program modules, or other data in a modulated data signal such as a carrier wave or other transport mechanism, and may include any information delivery media.
[0185] The preferred embodiments of the present application have been described above with reference to the accompanying drawings, but this does not limit the scope of the claims of the present application. Any modifications, equivalent substitutions, and improvements made by those skilled in the art without departing from the scope and substance of the embodiments of the present application shall be within the scope of the claims of the present application.
Claims
1. A method for rewriting tobacco business information queries based on direct preference optimization, characterized in that, The tobacco business information query rewriting method based on direct preference optimization includes the following steps: Construct a dataset of preferences for fuzzy questions in the tobacco business; Based on the aforementioned tobacco business fuzzy question preference dataset, a query rewriting task model is constructed; The query rewriting task model is trained using a direct preference optimization algorithm to obtain a query rewriting optimization model; Based on the query rewriting optimization model, the query rewriting of tobacco business information is completed.
2. The tobacco business information query rewriting method based on direct preference optimization according to claim 1, characterized in that, The construction of the tobacco business fuzzy question preference dataset includes the following steps: Construct the original fuzzy query list; Construct a high-quality list of subqueries; Construct a list of low-quality subqueries; Based on the original fuzzy query list, the high-quality subquery list, and the low-quality subquery list, a database triplet is obtained; Based on the database triples, a fuzzy question preference dataset for tobacco business is obtained.
3. The tobacco business information query rewriting method based on direct preference optimization according to claim 2, characterized in that, The construction of the original fuzzy query list includes the following steps: We obtain user-asked tobacco-related questions from public consultation platforms and filter out vague expressions to obtain user fuzzy query data. Simulated fuzzy questions are generated based on tobacco process specification information to obtain simulated question data; By combining a large language model with process document segmentation, fuzzy questions are automatically generated to obtain model-generated query data. Based on the user's fuzzy query data, the simulated question data, and the model-generated query data, an original fuzzy query list is obtained.
4. The tobacco business information query rewriting method based on direct preference optimization according to claim 2, characterized in that, The construction of a high-quality subquery list includes the following steps: Based on the original fuzzy query list, the information of the original fuzzy query is transformed into matching statement information that conforms to preset standard conditions by rewriting. Based on the original fuzzy query list, the large language model is guided by the instruction template to output structured subquery data; Based on the matching statement information and the structured subquery data, a list of high-quality subqueries is obtained.
5. The tobacco business information query rewriting method based on direct preference optimization according to claim 2, characterized in that, The construction of the low-quality subquery list includes the following steps: Based on the original fuzzy query list and the high-quality subquery list, statement information deviation data is generated using a large language model; Based on the deviation of the data from the stated information, a list of low-quality subqueries is obtained.
6. The tobacco business information query rewriting method based on direct preference optimization according to claim 1, characterized in that, The formula used to construct the query rewriting task model based on the tobacco business fuzzy question preference dataset includes: Where, q * The rewriting question generated by the query rewriting task model; Q is the fuzzy question preference dataset for the tobacco business; P θ Let θ be the conditional probability distribution of the candidate questions in the query rewriting task model under parameter θ; q0 is a fuzzy question.
7. The tobacco business information query rewriting method based on direct preference optimization according to claim 1, characterized in that, The process of training the query rewriting task model using the direct preference optimization algorithm to obtain the query rewriting optimization model includes the following steps: Initialize the model structure and parameters; Based on the query rewrite task model and the model structure and parameters, the tobacco business fuzzy question preference dataset is loaded to obtain the reference model and the current strategy model. Based on the reference model and the current policy model, the generation probability and the construction loss function are calculated using the direct preference optimization algorithm. Based on the probabilities and the constructed loss function, obtain the updated policy model; Based on the update strategy model, a query rewrite optimization model is obtained.
8. The tobacco business information query rewriting method based on direct preference optimization according to claim 7, characterized in that, The formulas used to calculate the generation probability and construct the loss function based on the direct preference optimization algorithm according to the reference model and the current policy model include: Among them, L DPO (π θ ;π ref ) represents the loss function; π θ For the current strategy model; π ref Here is the reference model; x is the user's original fuzzy query; y1 is the first rewritten candidate sentence; y2 is the second rewritten candidate sentence; y w Rewritten for higher quality; y l This is a poor rewrite; σ is the Sigmoid function; β is the temperature hyperparameter; π * For the optimal strategy model, E is the expected value; D is the data distribution.
9. A tobacco business information query and rewriting system based on direct preference optimization, characterized in that, It includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the tobacco business information query rewriting method based on direct preference optimization as described in any one of claims 1 to 8.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-executable instructions for causing a computer to perform the tobacco business information query rewriting method based on direct preference optimization as described in any one of claims 1 to 8.
Citation Information
Patent Citations
Training method and query method of query word rewriting model and related products
CN117743505A
Query information rewriting method and electronic equipment
CN118069917A
User medical inquiry quality improvement method and system based on preference optimization
CN118658635A
Reinforcement learning alignment model training method and system based on AI feedback
CN118735002A
Document retrieval query rewriting method and system based on large language model
CN120011482A