Query rewriting optimization method and device, electronic equipment and readable storage medium

By combining context-awareness and user feedback with a dynamic optimization mechanism, generative language models and rule bases work together to solve the transparency and adaptability problems of existing query rewriting technologies, achieving query rewriting with high accuracy and transparency.

CN121561084APending Publication Date: 2026-02-24THE PEOPLES INSURANCE CO (GRP) OF CHINA LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511984009.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-25
Publication Date
2026-02-24

AI Technical Summary

Technical Problem

Existing query rewriting technologies rely on manual rules or large language models, lacking transparency and dynamic optimization capabilities, making it difficult to effectively capture complex user intent, resulting in low query rewriting accuracy.

Method used

By combining context-aware interpretability generation with a dynamic optimization mechanism based on user feedback, a generative language model is used to generate candidate statements with rewriting types and reasons. The rewriting strategy is optimized by retrieving from a rule base, and user feedback updates the rule base and reinforcement learning model parameters.

Benefits of technology

It improved the accuracy and transparency of query rewriting, enabled continuous learning and adaptive optimization of the system, and enhanced user trust and interaction efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121561084A_ABST
    Figure CN121561084A_ABST
Patent Text Reader

Abstract

The invention provides a query rewriting optimization method and device, electronic equipment and a readable storage medium. The method comprises the steps of obtaining current query information input by a user and context information associated with the current query information; based on the current query information and the context information, generating at least one first candidate rewriting statement with a rewriting type and a rewriting reason through a generative language model, and retrieving through a rule base to obtain at least one second candidate rewriting statement; presenting the first candidate rewriting statement, the rewriting type and rewriting reason corresponding to the first candidate rewriting statement, and the second candidate rewriting statement on a user graphical interface; in response to a feedback operation of the user on the candidate rewriting statements, updating the rule base, and updating parameters of the reinforcement learning model; the reinforcement learning model is used for optimizing a generation strategy of the generative language model. In this way, by combining interpretability generation of context awareness and a dynamic optimization mechanism based on user feedback, the accuracy of query rewriting is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the fields of natural language processing and artificial intelligence, and in particular to a query rewriting optimization method, apparatus, electronic device, and readable storage medium. Background Technology

[0002] With the advancement of natural language processing technology, query rewriting has become a key technology for intelligent customer service, search engines, and dialogue systems. Its goal is to transform users' vague and brief original expressions into accurate and standardized queries in order to improve the accuracy of system understanding and interaction.

[0003] Currently, query rewriting mainly relies on template matching based on manually generated rules or end-to-end generation based on large language models. These existing solutions generally have the following limitations: First, the rewriting process is opaque, and users cannot understand the basis for the results, leading to insufficient trust in the system; second, they only support binary feedback such as "likes / dislikes," failing to effectively capture the complex intent of users who require adjustments based on partial satisfaction with the results; finally, the system's optimization capabilities are limited, typically only able to make static adjustments to existing queries, making it difficult to continuously learn from interactions and adapt to new query scenarios, resulting in low accuracy in query rewriting. Summary of the Invention

[0004] In view of this, embodiments of this application provide at least one query rewriting optimization method, apparatus, electronic device, and readable storage medium, which improves the accuracy of query rewriting by combining context-aware interpretability generation with a dynamic optimization mechanism based on user feedback.

[0005] This application mainly includes the following aspects: In a first aspect, embodiments of this application provide a query rewriting optimization method, the method comprising: Obtain the current query information input by the user and the context information associated with the current query information; Based on the current query information and the context information, at least one first candidate rewrite statement with rewrite type and rewrite reason is generated by generative language model, and at least one second candidate rewrite statement is obtained by rule base retrieval. The user graphical interface displays the first candidate rewrite statement, the rewrite type and rewrite reason corresponding to the first candidate rewrite statement, and the second candidate rewrite statement; In response to user feedback on the candidate rewritten statements, the rule base is updated, and the parameters of the reinforcement learning model are updated; the reinforcement learning model is used to optimize the generation strategy of the generative language model.

[0006] Secondly, embodiments of this application also provide a query rewriting optimization apparatus, the query rewriting optimization apparatus comprising: The information acquisition module is used to acquire the current query information input by the user and the context information associated with the current query information; The candidate generation module is used to generate at least one first candidate rewrite statement with rewrite type and rewrite reason based on the current query information and the context information through a generative language model, and to obtain at least one second candidate rewrite statement through rule base retrieval. An interactive presentation module is used to present the first candidate rewritten statement, the rewriting type and rewriting reason corresponding to the first candidate rewritten statement, and the second candidate rewritten statement on the user's graphical interface. The rewrite optimization module is used to update the rule base and update the parameters of the reinforcement learning model in response to the user's feedback operation on the candidate rewrite statement; the reinforcement learning model is used to optimize the generation strategy of the generative language model.

[0007] Thirdly, embodiments of this application also provide an electronic device, including: a processor, a memory, and a bus. The memory stores machine-readable instructions executable by the processor. When the electronic device is running, the processor and the memory communicate through the bus, and the machine-readable instructions are executed by the processor to perform the steps of the query rewrite optimization method as described above.

[0008] Fourthly, embodiments of this application also provide a computer-readable storage medium storing a computer program, which, when executed by a processor, performs the steps of the query rewrite optimization method as described above.

[0009] The query rewriting optimization method, apparatus, electronic device, and readable storage medium provided in this application acquire current query information input by the user and context information associated with the current query information; based on the current query information and context information, at least one first candidate rewritten statement with a rewriting type and rewriting reason is generated through a generative language model, and at least one second candidate rewritten statement is obtained through rule base retrieval; the first candidate rewritten statement, the corresponding rewriting type and rewriting reason, and the second candidate rewritten statement are presented on a user graphical interface; in response to user feedback on the candidate rewritten statements, the rule base is updated, and the parameters of the reinforcement learning model are updated; the reinforcement learning model is used to optimize the generation strategy of the generative language model. Thus, by combining context-aware interpretable generation with a dynamic optimization mechanism based on user feedback, the accuracy of query rewriting is improved.

[0010] To make the above-mentioned objectives, features and advantages of this application more apparent and understandable, preferred embodiments are described below in detail with reference to the accompanying drawings. Attached Figure Description

[0011] To more clearly illustrate the technical solutions of the embodiments of this application, the accompanying drawings used in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of this application and should not be regarded as a limitation of the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.

[0012] Figure 1 A flowchart of a query rewriting optimization method provided in an embodiment of this application is shown; Figure 2 This illustration shows one of the functional block diagrams of a query rewrite optimization apparatus provided in an embodiment of this application; Figure 3 This illustration shows a second functional block diagram of a query rewriting optimization apparatus provided in an embodiment of this application; Figure 4 A schematic diagram of the structure of an electronic device provided in an embodiment of this application is shown. Detailed Implementation

[0013] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. The components of the embodiments of this application described and shown in the accompanying drawings can generally be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of this application provided in the accompanying drawings is not intended to limit the scope of the claimed application, but merely represents selected embodiments of this application. All other embodiments obtained by those skilled in the art based on the embodiments of this application without inventive effort are within the scope of protection of this application.

[0014] To facilitate understanding of this application, the technical solutions provided in this application will be described in detail below with reference to specific embodiments.

[0015] Please see Figure 1 , Figure 1 This is a flowchart illustrating a query rewriting optimization method provided in an embodiment of this application. Figure 1 As shown in the embodiments of this application, the query rewriting optimization method includes the following steps: S101, obtain the current query information input by the user and the context information associated with the current query information.

[0016] Here, contextual information is primarily used to supplement and clarify the user's true intent in the current query. Typical sources include the user's conversation history and user profile data. The conversation history records recent, consecutive rounds of dialogue, helping the system understand which topic the current query continues or which part of previous information it is asking about. The user profile contains the user's attributes, preferences, and historical behavioral tags, providing prior information for the query's domain and specificity.

[0017] In this embodiment, taking an intelligent insurance customer service scenario as an example, a user might first inquire about "car insurance policy," then ask about "claims process," and finally input "how long will it take" for the current query. The system will collect the dialogue history from the most recent rounds (e.g., the last 3 rounds), including inquiries about "car insurance policy purchase" and "car insurance claims process." Simultaneously, it will obtain information from the user profile, such as the user's identity as a "private car owner," holding a "car insurance" policy, and their most recent claims record being a "scratch claim in May 2023." This contextual information collectively constitutes the semantic background for understanding the current query "how long will it take," meaning the user is actually concerned with the "processing cycle of car insurance claims." Before semantic encoding, Named Entity Recognition (NER) technology can be used to automatically identify entities with specific meanings from the current query information and contextual information, such as "car insurance," "claims," ​​and "private car owner," to focus on the core concepts.

[0018] S102, based on the current query information and the context information, at least one first candidate rewrite statement with rewrite type and rewrite reason is generated by generative language model, and at least one second candidate rewrite statement is obtained by rule base retrieval.

[0019] Here, this application employs a dual-channel parallel strategy to obtain rewriting candidates. The first candidate rewriting statement focuses on leveraging the semantic understanding and generation capabilities of the large model, combined with contextual information, to generate novel, relevant, and self-explanatory rewriting results. The second candidate rewriting statement focuses on reusing historical experience, quickly retrieving validated high-quality rewriting schemes from accumulated user feedback data.

[0020] In this embodiment, for the current query "How long does it take to process a claim?", the system constructs a structured prompt, inputting the current query, relevant entities retrieved from the context (such as "car insurance" and "review"), and explicit output format instructions (requiring the generation of a rewritten statement, type, and reason) into a generative language model (such as GPT-4). The model may output the first candidate rewritten statement "How long is the car insurance claim review period?", along with the rewriting type "semantic completion and synonym replacement" and the rewriting reason "Refer to the dialogue history to supplement the 'car insurance' scenario, and replace the colloquial 'how long' with the professional term 'review period'". On the other hand, the system simultaneously searches the rule base for historical queries semantically similar to "How long does it take to process a claim?" (such as "claim progress query"), and retrieves the rewritten statement that has been "confirmed" or "edited" most frequently by the user under that historical query, such as "car insurance claim processing time", as the second candidate rewritten statement.

[0021] S103, the first candidate rewrite statement, the rewrite type and rewrite reason corresponding to the first candidate rewrite statement, and the second candidate rewrite statement are presented in the user graphical interface.

[0022] The core purpose of this presentation is to make the thought process transparent and provide users with intuitive and convenient access to choices and feedback. The presentation not only displays multiple candidate rewritten statements, but more importantly, it shows the generation logic of the first candidate statement (rewrite type and rewrite reason) in an easy-to-understand way.

[0023] In this embodiment, the system merges the two candidate statements, "What is the review period for car insurance claims?" (with the rewriting type and reason) and "Car insurance claims processing time", into a single list, displayed below the user's original query, "How long does it take to process a claim?". In the list, the parts of each candidate statement that have changed compared to the original query (such as "car insurance claims", "review period", and "processing time") are highlighted in green. When the user hovers the mouse cursor over the highlighted "review period", the corresponding reason for the rewriting will be displayed nearby or in the sidebar: "Referring to the dialogue history and supplementing the 'car insurance' scenario, the colloquial 'how long' has been replaced with the professional term 'review period'." This presentation method allows users to clearly see "what has been changed" and "why it has been changed."

[0024] S104, in response to the user's feedback operation on the candidate rewritten statement, update the rule base and update the parameters of the reinforcement learning model; the reinforcement learning model is used to optimize the generation strategy of the generative language model.

[0025] Here, user feedback is the core driver of continuous system optimization. The system not only records feedback results, but more importantly, it transforms different types of feedback into quantifiable and learnable signals, which are used to update the short-term experience base (rule base) and the long-term policy model (reinforcement learning model), respectively. This allows the system to adjust recommendations in real time and gradually learn more general user preferences.

[0026] In this embodiment, assuming the user finds "review period" too technical and prefers the term "processing time," they perform an "edit" operation on the first candidate statement, changing it to "How long does it take to process car insurance claims" and clicking confirm. Upon receiving this "edited and confirmed" feedback, the system performs a two-way update: First, it stores the association between the current query "How long does it take to process a claim" and the user's final revised statement "How long does it take to process car insurance claims" as a new record in the rule base, assigning it an initial priority score for future quick retrieval. Second, it transforms this feedback event (original query, context, system-recommended action "review period," user's final action "processing time," and feedback type "edit") into a training sample for the reinforcement learning model. Specifically, the system encodes the query and context as a state (S), considers the system's original recommended "review period" as the action taken (A), and maps the "edit" operation to a moderately positive reward value (R = +0.7). Using this (S, A, R) data, the internal parameters of the reinforcement learning model are updated through temporal difference learning algorithms such as Deep Q Network (DQN). The goal is to enable the model to more accurately predict users' preferences for more colloquial expressions such as "processing time" when encountering similar situations in the future, thereby optimizing its recommendation strategy and indirectly affecting the prompt construction or result ranking of the generative language model.

[0027] Further, the step of generating at least one first candidate rewrite statement with rewrite type and rewrite reason based on the current query information and the context information through a generative language model includes: Step a1: Semantically encode the current query information and the context information to obtain vector representations of the current query information and the context information.

[0028] Here, the purpose of semantic encoding is to transform textual information into a mathematical form that can be processed and computed by a computer, laying the foundation for subsequent semantic matching and retrieval. The encoding process is typically accomplished using pre-trained language models (such as BERT and RoBERTa), which possess a deep understanding of the semantics of words and sentences. Encoding is divided into two levels: first, encoding the segmented text to obtain fine-grained word vectors, used to capture the semantics of core words; second, encoding complete sentences to obtain holistic sentence vectors, used to grasp the complete intent of the query or contextual fragments.

[0029] In this embodiment, the system encodes the current query "How long does it take to process a claim" and the statements extracted from the context (such as the historical dialogue "car insurance claim process" and the user profile "private car owner, car insurance user") using pre-trained models. For the query "How long does it take to process a claim", the system obtains its word vector set (WV-qs) after word segmentation and the sentence vector of the entire sentence (SV-qs). Similarly, for each piece of information in the context, the corresponding word vector set (such as WV-hs, WV-ps) and sentence vector set (such as SV-hs, SV-ps) are also calculated.

[0030] Step a2: Calculate similarity based on the vector representation and retrieve extended semantic association information from the context information.

[0031] Here, the retrieval of extended information aims to provide specific and relevant semantic material for the generation process, ensuring that the rewritten results closely adhere to the user's actual context. The retrieval process is achieved by calculating metrics such as cosine similarity between vectors, performed at both the word and sentence levels: at the word level, words semantically similar to the core words in the query are sought; at the sentence level, historical sentence fragments related to the overall intent of the query are searched. This guarantees that the retrieved information includes both replaceable synonyms and supplementary contextual logic.

[0032] In this embodiment, the system calculates the similarity between words such as "claims" and "how long" in the query word vector set (WV-qs) and all words in the context word vector sets (WV-hs, WV-ps), and selects the top-K most similar words. For example, "claims" is retrieved from historical dialogues, and "review" is retrieved from user profiles, forming a synonym extension set (SW). Simultaneously, the system calculates the similarity between the query sentence vector (SV-qs) and all sentence vectors in the context sentence vector sets (SV-hs, SV-ps), and finds the top-K most similar sentences. For example, it finds "the car insurance claims process generally requires submitting materials before entering the review stage" from historical dialogues, forming a related semantic extension set (SS). This extended information provides key semantic elements such as "car insurance" and "review" for generation.

[0033] Step a3: Based on the current query information, the semantic association extension information, and the predefined rewriting rules, construct structured prompt information.

[0034] Here, structured prompts are key instructions guiding generative language models to produce controllable and interpretable outputs. It's not a simple question, but rather a "program instruction manual" containing specific tasks, input materials, execution rules, and a strict output format. Predefined rewriting rules clarify the types of rewrites the model needs to generate (such as "synonym substitution" and "semantic completion") and their specific requirements, thus constraining the model's generation direction and ensuring the predictability and interpretability of the results.

[0035] In this embodiment, the system integrates the current query "How long does it take for a claim to be processed", the extended term set SW (e.g., "claims - claim / review; how long - cycle / duration"), the extended sentence set SS (e.g., "car insurance claim process...review stage"), and a preset rewriting rule template to construct detailed prompt information. This prompt information clearly defines the instruction model: "Based on the original question, a list of similar words, and examples of similar sentences, generate three types of rewriting results (synonym replacement, semantic completion, and comprehensive rewriting). Each type of result must include a complete sentence, a reason for rewriting, and a clear label of the rewriting type." This standardizes the generation task into a process of filling a structured table.

[0036] Below is a specific structured prompt message: """ Task: Based on the original question, list of similar words, and examples of similar sentences provided by the user, generate three types of rewriting results (synonym replacement rewriting, semantic completion rewriting, synonym replacement, and semantic completion rewriting). Each type of result must include a complete sentence, the reason for rewriting, and clearly indicate the rewriting type to ensure that it does not deviate from the core semantics of the original question.

[0037] # I. Reference information provided by users 1. Original question: {Query} 2. List of similar words: {SS} (Format: original core word - similar word 1 / similar word 2; e.g., "improve - enhance / strengthen; efficiency - effectiveness / rate") 3. Similar sentence examples: {SW} (Format instructions: 1. Sentence 1; 2. Sentence 2; 3. Sentence 3 (optional)) # II. Rewriting Execution Rules 1. Synonym replacement rewriting: Only words from the similar word list are used to replace the core words of the original question, without changing the sentence structure, question direction, or core semantics, ensuring complete equivalence of meaning.

[0038] 2. Semantic completion rewriting: Retain the core of the original question and add reasonable details (such as path, scene, scope, etc.) based on the logic of similar sentences. The added content must be related to the original question and no irrelevant information should be added.

[0039] 3. Synonym replacement and semantic completion rewriting: Combining rewriting rules 1 and 2.

[0040] # III. Output Format Requirements (Strictly follow this template for output; do not adjust the structure) ## Classification Rewriting Results ### (1) Rewriting type: Synonym replacement - Rewritten sentence: [Complete question after replacing the core word with a similar word] - Reason for rewriting: [Explain the specific words to be replaced, such as "replace 'original word 1' with 'similar word 1', replace 'original word 2' with 'similar word 2'"] ### (2) Rewriting type: semantic completion - Rewritten sentence: [Complete question with added details, such as "How to achieve this through XX method + core of the original question"] - Reason for rewriting: [Explain the additional details, such as "Referring to the 'XX logic' of similar sentence X, and adding 'XX details'"] ### (3) Rewriting types: semantic completion and synonym replacement - Rewritten sentence: [Complete question after replacing the core words with similar words, complete question after adding details] - Reason for rewriting: [Explain the specific words to be replaced and the additional details] """" Step a4: Input the structured prompt information into the generative language model to generate at least one first candidate rewrite statement with rewrite type and rewrite reason.

[0041] Here, after receiving strict structured prompts, the generative language model's free space is effectively constrained, and it instead executes a formatted generation according to the instructions. Its output is no longer free text, but a data block conforming to a predetermined structure, containing multiple fields (rewritten sentence, type, reason). This achieves synchronization between interpretability and the generation process, making the origin of each rewritten result clearly traceable.

[0042] In this embodiment, the constructed structured prompts are input into a generative language model such as GPT-4. The model will output formatted content according to the instructions, for example: Rewritten sentence: What is the review period for car insurance claims? Rewriting types: semantic completion and synonym replacement.

[0043] Reason for rewriting: Based on the dialogue history, a 'car insurance' scenario was added, and the colloquial 'how long' was replaced with the professional term 'review period'.

[0044] This output constitutes a complete first-choice rewrite statement with interpretable information. The system can parse this type of output to obtain one or more such structured results.

[0045] Furthermore, obtaining at least one second candidate rewrite statement through rule base retrieval includes: Step b1: Calculate the semantic similarity between the current query information and the historical query information in the rule base, and find similar historical query information based on the semantic similarity.

[0046] Here, the rule base stores a large number of historical user queries (historical query information) and their corresponding rewrite schemes. To quickly find the most relevant experiences to the current problem from this massive historical database, the system needs to quantify semantic similarity. This calculation is typically based on the vector representation technique described in step a1, evaluating the closeness of the current query's sentence vector to each historical query's sentence vector in the rule base by calculating metrics such as the cosine similarity. The system sets a similarity threshold or selects the top N most similar historical queries to determine which historical queries are sufficiently similar to the current query and can serve as valid references.

[0047] In this embodiment, the user's current input query is "How long will the claim take?". The system first calculates the sentence vector of this query, and then iterates through the sentence vectors of all historical queries in the rule base (such as "car insurance claim progress query", "insurance payment time", etc.), calculating the cosine similarity between each vector and the current query vector. Assuming that the calculation finds that "car insurance claim progress query" has the highest similarity to the current query and exceeds a preset threshold (such as 0.7), the system determines that the historical query is semantically similar to the current query and uses it as a similar search result.

[0048] Step b2: Obtain historical rewritten statements and corresponding priority data that are associated with the similar historical query information and verified by user feedback; the priority data is determined based on the feedback operations of historical users on the historical rewritten statements.

[0049] Here, after finding similar historical queries, the next step is to obtain the effective rewrite solutions and their quality scores accumulated under those historical queries and tested by actual user interaction. Each historical rewrite statement is associated with a dynamically updated priority data set. This data is not fixed but is calculated based on the feedback behavior of historical users towards that rewrite statement (such as confirmation, editing, and rejection). It quantifies the degree to which the rewrite statement is accepted and liked by the user group and is a core indicator of empirical reliability.

[0050] In this embodiment, the system retrieves all historically rewritten statements associated with the similar historical query "auto insurance claim progress query" found in step b1 from the rule base, such as "auto insurance claim processing time" and "auto insurance claim review period". Simultaneously, the system obtains the priority data corresponding to each rewritten statement. This priority data is calculated based on past user feedback on these statements: for example, the statement "auto insurance claim processing time" has historically been "confirmed" by 100 users, "edited" and adopted by 20 users, and "rejected" by 5 users. The system calculates its total score or average score according to a preset reward mapping rule (e.g., confirmation +1 point, editing +0.7 points, rejection -0.5 points), which serves as its current priority data. Therefore, "auto insurance claim processing time" may have a priority data of 85 points due to its high cumulative acceptance, while "auto insurance claim review period" may have a priority data of 70 points.

[0051] Step b3: Based on the priority data, select at least one of the highest priority statements from the acquired historical rewritten statements as the second candidate rewritten statements.

[0052] Here, after obtaining multiple candidate historical rewrite statements and their priority data, the system needs to make a final recommendation. The decision logic prioritizes the rewrite scheme with the best historical performance, i.e., the highest priority data. This reflects the system's trust in and reuse of collective experience, enabling it to quickly provide users with verified, high-quality answers, especially effective when handling common or repetitive queries.

[0053] In this embodiment, the system compares the priority data of multiple historical rewritten statements. It finds that "Auto Insurance Claim Processing Time" has the highest priority (85 points), while "Auto Insurance Claim Review Period" has a lower priority (70 points). Therefore, the system selects the highest-priority rewritten statement, "Auto Insurance Claim Processing Time," as the second candidate rewritten statement obtained in this search, ready to present it to the user.

[0054] Further, the step of presenting the first candidate rewritten statement, the rewritten type and rewritten reason corresponding to the first candidate rewritten statement, and the second candidate rewritten statement in the user graphical interface includes: Step c1: The first candidate rewrite statement and its corresponding rewrite type and rewrite reason are merged with the second candidate rewrite statement to form a candidate set.

[0055] The purpose of merging here is to integrate multiple candidate results from different generation strategies (innovative generation and empirical retrieval) to provide users with a unified view that can be compared and selected. The candidate set is a logical container that organizes all candidate rewritten statements to be presented and their associated information in an orderly manner, providing a convenient data structure for subsequent unified rendering and interactive processing.

[0056] In this embodiment, the system merges the first candidate rewritten statement "What is the car insurance claim review period?" generated by the generative language model, along with its associated rewritten type "semantic completion and synonym replacement" and rewritten reason "referencing dialogue history to supplement the 'car insurance' scenario, replacing the colloquial 'how long' with the professional term 'review period'", with the second candidate rewritten statement "car insurance claim processing time" retrieved from the rule base. The system may assign a temporary identifier or ranking weight to each candidate, ultimately forming a candidate set data structure containing two entries, ready to be passed to the front-end interface for rendering.

[0057] Step c2: The candidate set is presented in the user graphical interface; wherein, each candidate rewritten statement in the candidate set is compared with the current query information, and the changed text portion is highlighted.

[0058] The core of this presentation is the visualization of differences. Instead of simply listing candidate statements, the system proactively and intelligently compares each candidate with the original query, accurately identifying text segments that have been added, deleted, or replaced. These altered parts are then highlighted using techniques such as changing color, background color, or bolding. This allows users to instantly grasp the core modifications of each rewrite plan, significantly improving the efficiency and accuracy of information retrieval.

[0059] In this embodiment, the system renders the candidate set as a list in the user's graphical interface, displaying it below the user's original query "How long does it take to process a claim?". The system performs text difference comparison on each candidate statement: for "What is the review period for car insurance claims?", the system identifies the addition of "car insurance" and "review period", replacing "how long" with "how much"; for "Car insurance claim processing time", it identifies the addition of "car insurance" and "processing time". Then, when displayed on the interface, these identified changed parts (such as "car insurance", "review period", "processing time", etc.) are highlighted with a striking green background, while the remaining unchanged parts (such as "claims") maintain their normal style.

[0060] Step c3: In response to the user's triggering action on the identified text portion, display the corresponding reason for rewriting.

[0061] Here, this step allows for on-demand deepening of the explanation. Highlighting identifies what was changed, while further user interactions (such as mouse hover and clicks) reveal why those changes were made. By binding and associating specific reasons for the changes with the highlighted text, the system provides in-depth, context-sensitive explanations while maintaining a clean interface, satisfying users' needs for transparency and credibility.

[0062] In this embodiment, when a user questions the highlighted "review period" in the first candidate statement and hovers the mouse cursor over it, the system detects this hover trigger. Subsequently, the system searches for the rewriting reason associated with the highlighted candidate statement in step c1 and dynamically displays the corresponding detailed reason near the highlighted text (e.g., in the form of a tooltip): "Referring to the dialogue history to supplement the 'car insurance' scenario, the colloquial 'how long' has been replaced with the professional term 'review period'." This tooltip disappears after the user moves the mouse away. This method achieves elegant loading and unloading of explanatory information, resulting in a natural interaction and high information density.

[0063] Furthermore, the types of feedback operations include confirmation operations, editing operations, or rejection operations.

[0064] Here, the system provides users with three different granularity feedback types, aiming to capture the complete spectrum of user attitudes towards the rewritten results, ranging from complete satisfaction to complete dissatisfaction. This provides richer and more precise optimization signals than the traditional binary "like / dislike" feedback. The "confirm" action indicates that the user believes the candidate rewritten statement fully matches their intent and requires no modification; the "edit" action indicates that the user believes the candidate rewritten statement is partially or nearly correct, but requires manual text modification to fully meet expectations, providing specific instructions on "where to change" and "how to change it"; and the "reject" action indicates that the user believes the candidate rewritten statement completely fails to meet their intent.

[0065] In this embodiment, users can provide feedback on the presented candidate rewritten statements through buttons or similar controls on the interface. For example, for the candidate "What is the review period for car insurance claims?", if the user finds it accurate and professional, they can directly click the "Confirm" button, which represents a confirmation operation. If the user finds the term "review period" too formal and prefers to use "processing time", they can click the "Edit" button, modify the statement to "What is the processing time for car insurance claims?" in the pop-up text box, and then save it, which represents an editing operation. If the user believes that the rewritten statement completely deviates from the "insurance cancellation" related question they want to inquire about, they can click the "Reject" button, which represents a rejection operation. Through these three operations, the system accurately distinguishes between three different user states: "fully accepts", "partially accepts with adjustments", and "completely rejects", providing a key basis for subsequent differentiated optimization.

[0066] Further, update the rule base according to the following steps: Step d1: If the feedback operation is a confirmation operation, the candidate rewrite statement confirmed by the user is associated with the current query information and stored in the rule base, and the priority data of the confirmed candidate rewrite statement is increased.

[0067] Here, the confirmation action represents the user's direct affirmation of the system's recommendation. The system considers this a high-quality, successful interaction and needs to strengthen the influence of this experience in the rule base. Therefore, the system not only ensures that the "query-rewrite" correspondence is recorded in the database, but more importantly, by increasing its priority data (such as increasing its weight score), the rewritten statement has a higher probability of being retrieved and recommended to the user when encountering the same or similar queries in the future, thereby achieving rapid reuse of experience and reinforcement of high-quality results.

[0068] In this embodiment, the user confirms the candidate statement "car insurance claim processing time". The system then establishes or strengthens the mapping relationship between "how long does it take to process a claim" and "car insurance claim processing time" in the rule base. Simultaneously, based on preset reward rules (e.g., a +1.0 reward for confirmation), the system converts this reward value into a positive adjustment to the priority data of this record. For example, if its original priority score was 80 points, after this confirmation, the score is updated to 85 points, thus improving its ranking in future similar queries.

[0069] Step d2: If the feedback operation is an edit operation, the rewritten statement obtained after user editing is associated with the current query information and stored in the rule base, and initial priority data is assigned to the rewritten statement obtained after editing.

[0070] Here, editing is a crucial process where users provide "partially correct" feedback and contribute new knowledge. The system not only records user modifications but, more importantly, treats the revised statements as entirely new and valid experiences. This demonstrates the system's ability to learn new expressions and knowledge from user collaboration. An initial priority is assigned to this new statement, allowing it to enter the recommendation pool. Its priority is further adjusted based on feedback from other users (such as confirmation or rejection), thus achieving dynamic accumulation and optimization of experience.

[0071] In this embodiment, the user edits the candidate statement "What is the review period for car insurance claims?" to "What is the processing time for car insurance claims?". The system captures this final edited statement and treats it as a new "historical rewritten statement," associating it with the current query "How long does it take to process a claim?" and storing it in the rule base. Simultaneously, based on a preset reward for the editing operation (e.g., an editing operation reward of +0.7), the system assigns an initial priority score (e.g., 70 points) to this new record, making it a valid experience in the rule base for future retrieval and recommendation.

[0072] Step d3: If the feedback operation is a rejection operation, then the priority data of the candidate rewrite statements recommended in the rule base and associated with the current query information are reduced.

[0073] Here, a rejection is a user's negation of the system's recommendation. The system treats this as an interaction that requires learning. Its optimization logic is to reduce the priority of candidate rewritten statements actually recommended to the user in this round of dialogue. This means that for queries that lead to rejection, the system weakens the influence of currently relied-upon (potentially outdated or inaccurate) experience, creating opportunities for other potential, unrecommended rewritten solutions to be recommended in the future. This allows the rule base's recommendation strategy to adaptively adjust in response to negative user feedback.

[0074] In this embodiment, the user rejected all presented candidates (e.g., "car insurance claim review period" and "car insurance claim processing time"). The system determined this recommendation failure and identified entries in the rule base that were semantically related to the current query "how long does it take to process a claim" and were identical or highly similar to the rejected candidate statements. These entries were then negatively penalized in terms of priority data, for example, their priority score was reduced from 75 to 65. This causes the recommendation ranking of these rejected expressions to decrease the next time a similar query is encountered.

[0075] Further, the parameters of the reinforcement learning model are updated according to the following steps: Step e1: Determine the current query information and the context information as state features.

[0076] Here, state features are used to characterize the "environment" or "scenario" in which the system makes decisions. It needs to comprehensively and structurally reflect the user's current core needs (query information) and the background circumstances (contextual information) that influence the understanding of those needs. By encoding this textual information into a fixed-dimensional, machine-readable numerical vector (state features), the system can formally characterize the unique context of each interaction, providing a foundation for learning the strategy of "what action should be taken in what situation."

[0077] In this embodiment, for the current query "how long does the claim process take" and the associated contextual information (such as historical dialogues "car insurance claim process" and user profiles "private car owners"), the system performs semantic encoding and feature fusion. For example, it concatenates the query sentence vector with the vector of key contextual information to form a comprehensive, high-dimensional state feature vector. This state This uniquely represents the scenario of this interaction: "A car insurance user inquires about the processing time after consulting the claims process."

[0078] Step e2: The candidate rewritten statements are identified as action features.

[0079] Here, action features represent the "behaviors" or "choices" that the system can take in a given state. In this method, each candidate rewritten statement presented to the user is considered an "action" chosen by the system in that interaction. Each candidate rewritten statement is also encoded as an action feature vector. By characterizing actions, the reinforcement learning model can evaluate the expected value of different rewriting strategies (i.e., recommending different candidate statements) in a specific state.

[0080] In this embodiment, the system recommends the candidate rewritten statement "What is the car insurance claim review period?" to the user. This statement is independently encoded into an action feature vector. This action This represents the current state of the system. The specific recommendation decision made is as follows: We suggest that users use the professional term "review cycle".

[0081] Step e3: Determine a reward signal based on the type of the feedback operation; wherein the confirmation operation corresponds to a first reward value, the editing operation corresponds to a second reward value, the rejection operation corresponds to a third reward value, and the first reward value is greater than the second reward value, and the second reward value is greater than the third reward value.

[0082] Here, the reward signal is a direct, quantitative evaluation of the user's "action" quality by the system, serving as the core feedback driving model learning. By mapping feedback operations of different granularities to numerical rewards of varying magnitudes, a refined reward function is constructed: confirmation operations (complete satisfaction) receive the highest positive reward, encouraging the model to repeat such successful actions; editing operations (partial satisfaction) receive a moderate positive reward, affirming the approximate correctness of the direction while indicating room for improvement; and rejection operations (complete dissatisfaction) receive a negative reward (penalty), prompting the model to avoid making similar poor decisions in the future. This differentiated reward design is key to the model's ability to learn complex user preferences.

[0083] In this embodiment of the application, the user's actions recommended by the system The question "What is the review period for car insurance claims?" was edited to "What is the processing time for car insurance claims?". Based on preset reward mapping rules, such as +1.0 for confirmation, +0.7 for editing, and -0.5 for rejection, the system determines the immediate reward value for this round of interaction. This reward value quantifies the user's response to the action. The assessment is "partially accepted".

[0084] Step e4: Using the state features, the action features, and the reward signal, update the parameters of the reinforcement learning model through a temporal difference learning algorithm.

[0085] This is the core step in updating the parameters of a reinforcement learning model. The model uses collected "state-action-reward" empirical data and temporal difference learning algorithms (such as Q-learning and its deep version DQN) to update its internal policy or value function. The goal is to make the model's predictions of future rewards more accurate: specifically, to adjust the model parameters so that, for each state... Take action below The predicted value ( (value), which is closer to the immediate reward observed here. This is combined with the best expected future reward for the next state. Through numerous such iterative updates, the model gradually learns to evaluate the long-term value of all possible actions (different rewritten statements) for a given state (user query and context), thereby optimizing its recommendation strategy.

[0086] In this embodiment of the application, the system will record the current interaction experience (status). Action , award and the possible next state As a training sample, it is input into the depth In the network model, the model calculates the loss based on the temporal difference error and updates its network weights using the backpropagation algorithm. For example, its Value updates follow the formula: ,in It's the learning rate. It's a discount factor. After this update, when the model encounters a query scenario similar to state S in the future, for rewrite actions containing more technical terms like "approval period," its prediction will be... The value will be adjusted positively as needed; at the same time, the model may infer that in this scenario, more colloquial actions such as "processing time" may have higher potential value, thereby indirectly optimizing the generation or ranking strategy of the generative language model.

[0087] In one specific calculation embodiment, the system presets the reward values ​​as follows: confirmation operation +1.0, editing operation +0.7, and rejection operation -0.5; and sets the learning rate. =0.1, discount factor =0.95. Assuming the current state S, the original value corresponding to the action "Auto Insurance Claim Review Time" is... The value is 0.3, and the system receives feedback from the user regarding editing this result (reward). And assume the maximum Q value for the next state is 0.8. According to the formula Calculation, updated value Subsequently, the system, based on the updated... The recommendation priority of each candidate is recalculated, which improves the priority of "auto insurance claim review time" and makes it more likely to be recommended in subsequent similar queries.

[0088] This application provides a query rewriting optimization method, comprising: acquiring current query information input by a user and context information associated with the current query information; generating at least one first candidate rewriting statement with a rewriting type and a rewriting reason through a generative language model based on the current query information and the context information, and obtaining at least one second candidate rewriting statement through rule base retrieval; presenting the first candidate rewriting statement, the corresponding rewriting type and rewriting reason, and the second candidate rewriting statement on a user graphical interface; updating the rule base and updating the parameters of the reinforcement learning model in response to user feedback on the candidate rewriting statements; and using the reinforcement learning model to optimize the generation strategy of the generative language model. Thus, by combining context-aware interpretable generation with a dynamic optimization mechanism based on user feedback, the accuracy of query rewriting is improved.

[0089] Based on the same application concept, this application also provides a query rewriting optimization device corresponding to the query rewriting optimization method provided in the above embodiments. Since the principle of the device in this application is similar to the query rewriting optimization method in the above embodiments of this application, the implementation of the device can refer to the implementation of the method, and the repeated parts will not be described again.

[0090] Please see Figure 2 , Figure 2 This is one of the functional block diagrams of a query rewrite optimization device provided in an embodiment of this application. For example... Figure 2 As shown, the query rewrite optimization device 200 includes: The information acquisition module 201 is used to acquire the current query information input by the user and the context information associated with the current query information.

[0091] The candidate generation module 202 is used to generate at least one first candidate rewrite statement with rewrite type and rewrite reason through a generative language model based on the current query information and the context information, and to obtain at least one second candidate rewrite statement through rule base retrieval.

[0092] The interactive presentation module 203 is used to present the first candidate rewritten statement, the rewritten type and rewritten reason corresponding to the first candidate rewritten statement, and the second candidate rewritten statement on the user's graphical interface.

[0093] The rewrite optimization module 204 is used to update the rule base and update the parameters of the reinforcement learning model in response to the user's feedback operation on the candidate rewrite statement; the reinforcement learning model is used to optimize the generation strategy of the generative language model.

[0094] Furthermore, when the candidate generation module 202 generates at least one first candidate rewrite statement with rewrite type and rewrite reason based on the current query information and the context information using a generative language model, the candidate generation module 202 is specifically used for: Semantic encoding is performed on the current query information and the context information to obtain vector representations of the current query information and the context information; Similarity is calculated based on the vector representation, and extended information on semantic association is retrieved from the context information; Based on the current query information, the extended information of semantic association, and the predefined rewriting rules, a structured prompt message is constructed; The structured prompt information is input into the generative language model to generate at least one first candidate rewrite statement with rewrite type and rewrite reason.

[0095] Furthermore, when the candidate generation module 202 is used to obtain at least one second candidate rewritten statement through rule base retrieval, the candidate generation module 202 is specifically used for: Calculate the semantic similarity between the current query information and the historical query information in the rule base, and find similar historical query information based on the semantic similarity; Obtain historical rewritten statements and their corresponding priority data that are associated with the similar historical query information and verified by user feedback; the priority data is determined based on the feedback operations of historical users on the historical rewritten statements. Based on the priority data, at least one of the highest priority statements is selected from the acquired historical rewritten statements as the second candidate rewritten statements.

[0096] Furthermore, when the interactive presentation module 203 is used to present the first candidate rewritten statement, the rewritten type and rewritten reason corresponding to the first candidate rewritten statement, and the second candidate rewritten statement on the user's graphical interface, the interactive presentation module 203 is specifically used for: The first candidate rewrite statement and its corresponding rewrite type and rewrite reason are merged with the second candidate rewrite statement to form a candidate set; The candidate set is presented in the user graphical interface; wherein, each candidate rewritten statement in the candidate set is compared with the current query information, and the changed text portion is highlighted. In response to a user's action on the identified text section, display the corresponding reason for the rewrite.

[0097] Further, please refer to Figure 3 , Figure 3 This is a second functional block diagram of a query rewrite optimization device provided in an embodiment of this application. For example... Figure 3 As shown, the query rewrite optimization device 200 also includes: The first update module 205 is used to, if the feedback operation is a confirmation operation, associate the candidate rewrite statement confirmed by the user with the current query information and store it in the rule base, and increase the priority data of the confirmed candidate rewrite statement.

[0098] The second update module 206 is used to, if the feedback operation is an edit operation, associate the rewritten statement obtained after user editing with the current query information and store it in the rule base, and assign initial priority data to the rewritten statement obtained after editing.

[0099] The third update module 207 is used to reduce the priority data of the candidate rewrite statements that are recommended in the rule base and associated with the current query information if the feedback operation is a rejection operation.

[0100] Furthermore, such as Figure 3 As shown, the query rewrite optimization device 200 also includes: The first determining module 208 is used to determine the current query information and the context information as state features.

[0101] The second determining module 209 is used to determine the candidate rewritten statement as an action feature.

[0102] The third determining module 210 is used to determine a reward signal based on the type of the feedback operation; wherein the confirmation operation corresponds to a first reward value, the editing operation corresponds to a second reward value, the rejection operation corresponds to a third reward value, and the first reward value is greater than the second reward value, and the second reward value is greater than the third reward value.

[0103] The reinforcement learning module 211 is used to update the parameters of the reinforcement learning model using the state features, the action features, and the reward signal through a temporal difference learning algorithm.

[0104] This application provides a query rewriting optimization device, comprising: an information acquisition module for acquiring current query information input by a user and context information associated with the current query information; a candidate generation module for generating at least one first candidate rewritten statement with a rewriting type and a rewriting reason based on the current query information and context information using a generative language model, and obtaining at least one second candidate rewritten statement through rule base retrieval; an interactive presentation module for presenting the first candidate rewritten statement, the corresponding rewriting type and rewriting reason, and the second candidate rewritten statement on a user graphical interface; and a rewriting optimization module for updating the rule base and updating the parameters of a reinforcement learning model in response to user feedback on the candidate rewritten statements; the reinforcement learning model is used to optimize the generation strategy of the generative language model. Thus, by combining context-aware interpretable generation with a dynamic optimization mechanism based on user feedback, the accuracy of query rewriting is improved.

[0105] Based on the same application concept, please refer to Figure 4 , Figure 4 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Figure 4 As shown, the electronic device 400 includes a processor 410, a memory 420, and a bus 430.

[0106] The memory 420 stores machine-readable instructions executable by the processor 410. When the electronic device 400 is running, the processor 410 and the memory 420 communicate through the bus 430. When the machine-readable instructions are run by the processor 410, the steps of the query-rewrite optimization method provided in the above embodiment are executed. For specific implementation, please refer to the method embodiment, which will not be repeated here.

[0107] Based on the same concept, this application also provides a computer-readable storage medium storing a computer program. When the computer program is run by a processor, it executes the steps of the query rewrite optimization method provided in the above embodiments. For specific implementation details, please refer to the method embodiments, which will not be repeated here.

[0108] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working process of the above-described apparatus and unit can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.

[0109] In the embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. The apparatus embodiments described above are merely illustrative. For example, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. Furthermore, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Additionally, the displayed or discussed mutual couplings, direct couplings, or communication connections may be through some communication interfaces; indirect couplings or communication connections between devices or units may be electrical, mechanical, or other forms.

[0110] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0111] In addition, the functional units in the embodiments provided in this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.

[0112] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0113] It should be noted that similar labels and letters in the following figures indicate similar items. Therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures. In addition, the terms "first", "second", "third", etc. are used only to distinguish descriptions and should not be construed as indicating or implying relative importance.

[0114] Finally, it should be noted that the above-described embodiments are merely specific implementations of this application, used to illustrate the technical solutions of this application, and not to limit them. The protection scope of this application is not limited thereto. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that any person skilled in the art can still modify or easily conceive of changes to the technical solutions described in the foregoing embodiments, or make equivalent substitutions for some of the technical features, within the scope of the technology disclosed in this application; and these modifications, changes, or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application. All should be covered within the protection scope of this application. Therefore, the protection scope of this application should be determined by the protection scope of the claims.

Claims

1. A query rewriting optimization method, characterized in that, The method includes: Obtain the current query information input by the user and the context information associated with the current query information; Based on the current query information and the context information, at least one first candidate rewrite statement with rewrite type and rewrite reason is generated by generative language model, and at least one second candidate rewrite statement is obtained by rule base retrieval. The user graphical interface displays the first candidate rewrite statement, the rewrite type and rewrite reason corresponding to the first candidate rewrite statement, and the second candidate rewrite statement; In response to user feedback on the candidate rewritten statements, the rule base is updated, and the parameters of the reinforcement learning model are updated; the reinforcement learning model is used to optimize the generation strategy of the generative language model.

2. The query rewriting optimization method according to claim 1, characterized in that, The step of generating at least one first candidate rewrite statement with rewrite type and rewrite reason based on the current query information and the context information through a generative language model includes: Semantic encoding is performed on the current query information and the context information to obtain vector representations of the current query information and the context information; Similarity is calculated based on the vector representation, and extended information on semantic association is retrieved from the context information; Based on the current query information, the extended information of semantic association, and the predefined rewriting rules, a structured prompt message is constructed; The structured prompt information is input into the generative language model to generate at least one first candidate rewrite statement with rewrite type and rewrite reason.

3. The query rewriting optimization method according to claim 1, characterized in that, The step of obtaining at least one second candidate rewrite statement through rule base retrieval includes: Calculate the semantic similarity between the current query information and the historical query information in the rule base, and find similar historical query information based on the semantic similarity; Obtain historical rewritten statements and their corresponding priority data that are associated with the similar historical query information and verified by user feedback; the priority data is determined based on the feedback operations of historical users on the historical rewritten statements. Based on the priority data, at least one of the highest priority statements is selected from the acquired historical rewritten statements as the second candidate rewritten statements.

4. The query rewriting optimization method according to claim 1, characterized in that, The step of presenting the first candidate rewritten statement, the rewritten type and rewritten reason corresponding to the first candidate rewritten statement, and the second candidate rewritten statement in the user graphical interface includes: The first candidate rewrite statement and its corresponding rewrite type and rewrite reason are merged with the second candidate rewrite statement to form a candidate set; The candidate set is presented in the user graphical interface; wherein, each candidate rewritten statement in the candidate set is compared with the current query information, and the changed text portion is highlighted. In response to a user's action on the identified text section, display the corresponding reason for the rewrite.

5. The query rewriting optimization method according to claim 3, characterized in that, The types of feedback operations include confirmation, editing, or rejection.

6. The query rewriting optimization method according to claim 5, characterized in that, Update the rule base according to the following steps: If the feedback operation is a confirmation operation, the candidate rewrite statement confirmed by the user is associated with the current query information and stored in the rule base, and the priority data of the confirmed candidate rewrite statement is increased. If the feedback operation is an edit operation, the rewritten statement obtained after the user edits is associated with the current query information and stored in the rule base, and initial priority data is assigned to the rewritten statement obtained after editing; If the feedback operation is a rejection operation, then the priority data of the candidate rewrite statements recommended in the rule base and associated with the current query information are reduced.

7. The query rewriting optimization method according to claim 6, characterized in that, Update the parameters of the reinforcement learning model according to the following steps: The current query information and the context information are determined as state features; The candidate rewritten statements are identified as action features; A reward signal is determined based on the type of feedback operation; wherein the confirmation operation corresponds to a first reward value, the editing operation corresponds to a second reward value, and the rejection operation corresponds to a third reward value, and the first reward value is greater than the second reward value, and the second reward value is greater than the third reward value; The parameters of the reinforcement learning model are updated using the state features, the action features, and the reward signal through a temporal difference learning algorithm.

8. A query rewrite optimization device, characterized in that, The query rewriting optimization device includes: The information acquisition module is used to acquire the current query information input by the user and the context information associated with the current query information; The candidate generation module is used to generate at least one first candidate rewrite statement with rewrite type and rewrite reason based on the current query information and the context information through a generative language model, and to obtain at least one second candidate rewrite statement through rule base retrieval. An interactive presentation module is used to present the first candidate rewritten statement, the rewriting type and rewriting reason corresponding to the first candidate rewritten statement, and the second candidate rewritten statement on the user's graphical interface. The rewrite optimization module is used to update the rule base and update the parameters of the reinforcement learning model in response to the user's feedback operation on the candidate rewrite statement; the reinforcement learning model is used to optimize the generation strategy of the generative language model.

9. An electronic device, characterized in that, include: The device includes a processor, a memory, and a bus. The memory stores machine-readable instructions executable by the processor. When the electronic device is running, the processor communicates with the memory via the bus. The machine-readable instructions are executed by the processor to perform the steps of the query rewrite optimization method as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, performs the steps of the query rewrite optimization method as described in any one of claims 1 to 7.