Financial intelligent question and answer method and system based on mixed retrieval and dynamic query

By building a FAQ knowledge base and a structured database of the Finance Department, combined with multimodal intention recognition and dynamic database query, the problems of poor retrieval results and low personalized reply efficiency in the existing financial question-and-answer system are solved, and efficient and accurate financial intelligent question-and-answer system are achieved.

CN120256574APending Publication Date: 2025-07-04CHONGQING UNIV OF POSTS & TELECOMM
View PDF 0 Cites 13 Cited by

Patent Information

Application Number
CN202510339135.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-21
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

The existing financial question and answer system is difficult to take into account both the proper noun and context semantics, resulting in poor low-frequency word search results, ambiguity of polysemes, and lack of dynamic database query logic, making it impossible to generate accurate personalized replies, which are inefficient.

Method used

Using a financial intelligent question-and-answer method based on hybrid search and dynamic query, the FAQ knowledge base and the Finance Department structured database are constructed, combined with multimodal intention recognition, semantic rewriting and dynamic database query, a large language model is used for intent recognition and semantic rewriting, standardized query statements are generated, and candidate lists are obtained through the mixed search mode, and the personalized problem is judged and the structured database is called to generate query results.

Benefits of technology

It improves the recall and accuracy of searches, can generate accurate and personalized replies, and improves the efficiency and accuracy of financial questions and answers.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120256574A_ABST
    Figure CN120256574A_ABST
Patent Text Reader

Abstract

The invention belongs to the field of intelligent question-answering systems, and relates to a financial intelligent question-answering method and system based on mixed retrieval and dynamic query, and the method comprises the steps: constructing an FAQ library and a financial office database; receiving a user question, calling a historical chat record of the user, and calling the large language model to perform intention recognition; calling a large language model to perform semantic rewriting in combination with the real intention of the user to generate a standardized query statement; obtaining a candidate list from the FAQ library by adopting a mixed retrieval mode; if no candidate questions and answers exist in the candidate list, replying the standardized answers; otherwise, screening candidate questions and answers of Top-N, judging whether the candidate questions and answers of Top-1 are personalized questions or not, if yes, calling a financial office database to generate a query result, and replying according to the query result and the candidate questions and answers of Top-N; according to the method, semantic retrieval and keyword retrieval are mixed through the user intention classification result and the conflict detection rule, the candidate list is obtained according to the mixed retrieval mode, and the retrieval recall rate is increased.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of natural language processing and intelligent question-answering systems, and relates to a financial intelligent question-answering method and system based on hybrid retrieval and dynamic query. Background Art

[0002] Nowadays, financial knowledge question-answering systems can help enterprise employees quickly obtain the required financial information in their daily work, improving work efficiency and quality. At present, due to the huge and diverse amount of financial knowledge information, traditional retrieval methods are difficult to meet the needs of employees. Recently, financial question-answering systems based on large models have become a research hotspot.

[0003] Existing financial question-answering systems are mostly based on keyword matching or single-modal processing technologies. For example, traditional OCR technology can only extract text information in pictures, but cannot recognize and label unstructured content such as context semantics and symbols. In addition, existing retrieval algorithms mostly adopt single modes such as BM25 or semantic vectors, making it difficult to take into account proper nouns and context language, resulting in problems such as poor retrieval effects for low-frequency words and polysemy ambiguities. At the same time, existing systems lack dynamic database query logic and cannot generate accurate responses for users' personalized financial data (such as tuition fees, reimbursement amounts), relying on manual secondary confirmation, with low efficiency. Summary of the Invention

[0004] To solve the above problems of the existing technology, the present invention adopts a financial intelligent question-answering method based on hybrid retrieval and dynamic query, including:

[0005] S1. Construct an FAQ knowledge base and a structured database of the financial department;

[0006] S2. Receive the question currently input by the user, retrieve the user's historical chat records, and call a large language model to perform intent recognition on the question currently input by the user and the historical chat records to obtain the user's true intent and the user intent classification result;

[0007] S3. Call a large language model to semantically rewrite the question currently input by the user in combination with the user's true intent to generate a standardized query statement;

[0008] S4. Obtain a candidate list from the FAQ knowledge base in a hybrid retrieval mode according to the standardized query statement and the user intent classification result; if there is no candidate question and answer in the candidate list, return a standardized answer; otherwise, execute step S5;

[0009] S5. Screen the Top-N candidate Q&As in the candidate list, and determine whether the Top-1 candidate Q&A is a personalized question. If so, call the structured database of the finance department to generate a query result based on the Top-1 candidate Q&A, and reply according to the query result and the Top-N candidate Q&As; otherwise, directly reply according to the Top-N candidate Q&As; where N is the screening threshold.

[0010] On the other hand, the present invention adopts a financial intelligent Q&A system based on hybrid retrieval and dynamic query. This system is used to execute the above-mentioned financial intelligent Q&A method based on hybrid retrieval and dynamic query, including:

[0011] A multimodal intention recognition module, which is used to call a large language model to recognize the intention of the question currently input by the user and the historical chat records.

[0012] A semantic rewriting module, which is used to call a large language model to semantically rewrite the question input by the user in combination with the true intention of the user to generate a standardized query statement.

[0013] A hybrid retrieval and re-ranking module, which is used to obtain a candidate list from the FAQ knowledge base in a hybrid retrieval mode according to the standardized query statement and the user intention classification result; if there is no candidate Q&A in the candidate list, reply with a standardized answer; otherwise, screen the Top-N candidate Q&As in the candidate list.

[0014] A dynamic database query module: determine whether the Top-1 candidate Q&A is a personalized question. If so, call the structured database of the finance department to generate a query result.

[0015] A reply module: call a large model to generate a reply using the query result, the Top-N candidate Q&As, and a pre-set Prompt template.

[0016] Beneficial effects:

[0017] 1. The present invention combines semantic retrieval and keyword retrieval through the user intention classification result and conflict detection rules, and obtains a candidate list according to the mixed retrieval mode, improving the recall rate of retrieval; 2. When the present invention combines the semantic matching similarity and the keyword matching similarity by weighting, a dynamic threshold adjustment item is added to adapt to the complex scenarios of multi-intention superposition and non-linear relationship, thereby improving the recall rate of retrieval; 3. The present invention uses a priority ranking table of proprietary nouns in the financial field to select the proprietary nouns with high priority in conflicts, which is more in line with the actual Q&A situation and improves the accuracy of Q&A; 4. The present invention uses fixed Python code to detect whether the database query statement generated by the large model is complete. If it is incomplete, the large model is called to regenerate the database query statement according to the error information of the fixed Python code and the query keywords, improving the retrieval efficiency and accuracy; 5. The present invention determines whether the top-1 candidate Q&A is a personalized question. If so, the structured database of the finance department is called to generate a query result according to the top-1 candidate Q&A, so as to generate a precise response for the user's personalized financial data and improve the personalized response efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] Figure 1 It is a flowchart of a financial intelligent Q&A method based on hybrid retrieval and dynamic query provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0019] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0020] As Figure 1 shown, on the one hand, the present invention adopts a financial intelligent Q&A method based on hybrid retrieval and dynamic query, including:

[0021] S1. Construct a FAQ knowledge base and a structured database of the finance department;

[0022] Constructing the FAQ knowledge base includes: obtaining the conversation text between users and consultants in the financial field; extracting questions and answers matching the questions from the conversation text to obtain multiple Q&A pairs.

[0023] S2. Receive the question currently input by the user, retrieve the historical chat records of the user within 24 hours, and call the large language model to perform intention recognition on the question currently input by the user and the historical chat records to obtain the true intention of the user and the user intention classification result;

[0024] The intention recognition of the user's current input question and historical chat records includes: determining whether the user's current input question contains a picture. If so, input the user's current input question (including text + picture) into a multimodal large model, and let the multimodal large model initially recognize the accurate content of the picture according to the text information input by the user to obtain structured semantic information; input the structured semantic information and the user's historical chat records within 24 hours into a large language model to recognize the user's true intention and the user intention classification result; otherwise, directly input the user's current input question and the user's historical chat records within 24 hours into the large language model to recognize the user's true intention and the user intention classification result.

[0025] In one embodiment, the multimodal large model is qwen-vl-max, and the large language model is DeepSeeK-R1.

[0026] Traditional OCR (Optical Character Recognition) technology only simply recognizes and extracts text, without considering picture frames, structures, and semantic information. Directly inputting it into a large language model for intention recognition may lead to biased intentions or even completely opposite intentions. However, the multimodal large model can recognize picture frames, structures, and semantic information, and will not have problems such as unsmooth sentences or opposite meanings, and can more accurately recognize picture information.

[0027] Example:

[0028] The user sent a picture of a travel expense bill form from Chongqing to Beijing, which contains a list of transportation expenses, accommodation expenses, food expenses, etc. At the same time, the user also sent a text "How much can be reimbursed?"

[0029] At this time, the multimodal large model will output two parts:

[0030] (1) It will recognize all the content information on the picture, including all information such as numbers, amounts, names, locations, etc., and will automatically optimize the layout and output.

[0031] (2) According to the user's question, it will output that the possible intention of the user is: the user wants to inquire about business trips outside the city (cities other than Chongqing), and whether transportation expenses (xx yuan), accommodation expenses (xx yuan), and food expenses (xx yuan) can be reimbursed.

[0032] By fine-tuning the large language model with a dataset with intention labels, the large language model can learn to classify the user's questions into predefined intention categories. In one embodiment, as shown in Table 1, the intention classification is a total of 5 categories; the user intention classification result is a probability distribution vector: p = [p1, p2,..., p5], where p j represents the probability that the user's intention belongs to the jth type of intention.

[0033] Table 1 Intention classification categories:

[0034]

[0035]

[0036] S3. Call the large language model to semantically rewrite the question input by the user in combination with the user's true intention, and generate a standardized query statement;

[0037] Some users have literal problems such as typos and incomplete semantics. Moreover, if only one or two sentences of the user's original text are used for information retrieval, the retrieved information is likely to be inaccurate, and then the intention recognition will be meaningless. Let the large language model rewrite the user's question according to the user's true intention, which can improve the subsequent retrieval accuracy; among them, the large language model can be DeepSeeK-R1.

[0038] S4. According to the standardized query statement and the user intention classification result, adopt a hybrid retrieval mode to obtain a candidate list from the FAQ knowledge base; if there is no candidate Q&A in the candidate list, return the standardized answer; otherwise, execute step S5; where the FAQ knowledge base is a frequently asked questions and answers knowledge base, which contains rich Q&A content;

[0039] The standardized answer is "I can only answer your questions related to finance. Please re-enter your question in detail!"

[0040] Keyword retrieval can make up for the disadvantage that semantic retrieval has poor retrieval effect on abbreviations, proper nouns (such as financial professional terms) or low-frequency words. Semantic retrieval can make up for the disadvantage that keyword retrieval cannot understand the context or polysemous words (for example, "apple" cannot distinguish between a company name and a fruit).

[0041] Parallelly obtaining the candidate list from the FAQ knowledge base includes:

[0042] S41. Set the threshold Q = 0.6, and calculate the semantic matching similarity S between the standardized query statement and the Q&A in the FAQ knowledge base respectively by using semantic retrieval and keyword retrieval i 1 and the keyword matching similarity S i 2 ; where i is the index of the Q&A in the FAQ knowledge base;

[0043] In the test set, as shown in Table 2, when the threshold Q = 0.6, the retrieval recall rate reaches 85%, and the precision rate reaches 78%; compared with Q = 0.5 and Q = 0.7, the recall rate has increased, but the precision rate has decreased (the noise has increased). To balance the two evaluation indicators, the threshold Q = 0.6 is selected.

[0044] Table 2 Comparison of Threshold Selection

[0045] Threshold Q Recall rate Precision rate 0.5 90% 66% 0.6 85% 78% 0.7 72% 82%

[0046] S42. Select the Q&A in the FAQ knowledge base whose keyword matching similarity with the standardized query statement is not less than the threshold Q as the candidate Q&A for keyword retrieval, and sort the candidate Q&A for keyword retrieval according to the keyword matching similarity to obtain the candidate list L2 for keyword retrieval;

[0047] S43. Use the user intention classification result to perform weighted combination of the semantic matching similarity and the keyword matching similarity of the Q&A in the FAQ knowledge base to obtain the final matching similarity of the Q&A in the FAQ knowledge base

[0048] The final matching similarity between the standardized query statement and each Q&A in the FAQ knowledge base is:

[0049]

[0050] where J is the number of intention categories, α is the keyword retrieval weight, β is the semantic retrieval weight, and α + β = 1, and the value is determined by the intention classification result. is the dynamic threshold adjustment term, σ() is the Sigmoid function, k is the adjustment coefficient, and the default value of k is 5, which controls the steepness of the weight; are the basic weights of each intention category j for keyword retrieval and semantic retrieval respectively. The basic weights are defined according to the dependence degree of each intention category on keyword retrieval (BM25) and semantic retrieval.

[0051] In the financial Q&A scenario, user questions often contain multiple intention superpositions (such as involving policy terms and personalized numerical queries at the same time) or non-linear semantic relationships (such as fuzzy expressions that need to be understood in combination with the context). Traditional linear weighting cannot adapt to such complex requirements and may overly favor a certain type of retrieval method. Therefore, the present invention introduces a dynamic threshold adjustment term to adapt to the complex scenarios of multiple intention superpositions and non-linear relationships. Specifically, the dynamic threshold adjustment term maps the intention probability to a non-linear adjustment factor, smoothly scales the independent weight of each intention, and can ensure that high-probability intentions dominate the retrieval direction, and low-probability intentions contribute auxiliary weights, so as to cover the multi-intention association results. And α emphasizes the enhancement of keyword retrieval by high-confidence intentions, and β balances the dependence of the remaining intentions on semantic retrieval through 1 - p j to avoid a single intention monopolizing the weight. j balances the dependence of the remaining intentions on semantic retrieval and avoids a single intention monopolizing the weight.

[0052] The non - linear characteristics of the sigmoid function enable the weights to rapidly approach the upper limit when the probability of a certain type of intention exceeds 70% (high confidence), strengthening the priority of the corresponding retrieval pattern; when the intention probability is lower than 30% (low confidence), the weights decay exponentially to suppress the interference of irrelevant noise; within the fuzzy interval of 30% - 70%, the weights change continuously to adapt to the scenario of intention uncertainty.

[0053] In one embodiment, the basic weights are defined As shown in Table 3.

[0054] Table 3 Basic weights of each type of intention for keyword retrieval and semantic retrieval

[0055] Intention category Basic weight of keyword retrieval Basic weight of semantic retrieval Policy clause query 0.7 0.3 Personalized numerical query 0.2 0.8 Operation process consultation 0.5 0.5 Status tracking query 0.6 0.4 Concept explanation inquiry 0.3 0.7

[0056] In one embodiment, the algorithm for vector - based semantic retrieval is obtained by calculating the cosine similarity, and the retrieval algorithm based on keywords is obtained by the BM25 algorithm.

[0057] S44. Use the matching similarity between the FAQ knowledge base and the standardized query statement The Q&A with a matching similarity not less than the threshold Q as the candidate Q&A for hybrid retrieval, and sort the candidate Q&A for hybrid retrieval according to the matching similarity to obtain the candidate list L3 for hybrid retrieval;

[0058] S45. Use the conflict detection rules to adjust the candidate list L3 for hybrid retrieval according to the candidate list L2 for keyword retrieval to obtain the final candidate list.

[0059] Using the conflict detection rules to adjust the candidate list L3 for hybrid retrieval according to the candidate list L2 for keyword retrieval includes:

[0060] S451: Construct a proprietary noun library for the financial field and a priority ranking table for proprietary nouns in the financial field;

[0061] Data sources: university financial policy documents (such as "Measures for the Management of Funds", "Detailed Rules for Reimbursement"), high - frequency entities in historical user questions (such as "JZ - 2023 project", "Horizontal project number"), field names in the structured database of the finance department (such as "Tuition amount", "Date of labor service payment").

[0062] Storage structure of the proprietary noun library for the financial field:

[0063]

[0064] The priority ranking table for proprietary nouns in the financial field is shown in Table 4;

[0065] Table 4 Priority ranking table for proprietary nouns in the financial field

[0066]

[0067] S452: For the candidate lists L2 and L3 of keyword retrieval and hybrid retrieval respectively, extract proper nouns according to the proper noun library in the financial field to obtain the proper noun sets of the candidate lists L2 and L3.

[0068] Segment the Q&A FAQs in the candidate list L2 of keyword retrieval, match the entities in the proper noun library, and obtain the proper noun set included in each FAQ in the candidate list L2 (for example, FAQ1 contains the "JZ-2023 project"). According to the proper noun sets of all FAQs in the candidate list L2, obtain the proper noun set corresponding to the candidate list L2.

[0069] Similarly, the proper noun set of the candidate list L3 can be obtained.

[0070] S453: Combine the proper nouns that belong to the proper noun set of the candidate list L2 but not to the proper noun set of the candidate list L3 to obtain the conflicting proper noun set E = {e1, e2,..., e M}, and select the proper noun with the highest priority in the proper noun set E according to the proper noun priority sorting table in the financial field Among all the candidate Q&As in the candidate list L2 that contain the proper noun select the candidate Q&A with the highest keyword matching similarity and insert it into the Top-3 position in the candidate list L3 (the original Top-3 positions are shifted backward) to obtain the final candidate list; where e m is the proper noun in the proper noun set E, and M is the number of conflicting proper nouns.

[0071] Example scenario:

[0072] User's question: "How to query the funds of the JZ-2023 project?"

[0073] The candidate list of keyword retrieval contains the FAQ: "To query the JZ-2023 project, you need to log in to the financial system."

[0074] The candidate list of hybrid retrieval includes the FAQ: "The query process for the funds of all projects is similar. Please log in to the system."

[0075] The hybrid retrieval does not mention the specific project name "JZ-2023", and may not answer the question.

[0076] Therefore, adjust the candidate list of hybrid retrieval through the candidate list of keyword retrieval, and insert "To query the JZ-2023 project, you need to log in to the financial system." in the candidate list of keyword retrieval into the top 3 of the candidate list of hybrid retrieval.

[0077] When the final candidate list does not contain candidate Q&As, it indicates that the standardized query statement does not match the preset FAQ knowledge base. It may be due to issues such as unclear user questions, irrelevance to finance, involvement in sensitive topics, or FAQ question settings. At this time, a standardized response is returned: "I can only answer your finance-related questions. Please re-enter your question in detail!"

[0078] In one embodiment, a conflict word list in the financial field is preset (for example, "reimbursement" corresponds to synonyms such as "offsetting accounts" and "write-off"). When the user mentions certain professional terms, the corresponding preset conflict table in the financial field is automatically matched and then retrieved one by one, and the different expressions in the two sets of retrieval results are automatically aligned.

[0079] Specifically, it includes:

[0080] Step 1: Construct a conflict word list in the financial field

[0081]

[0082]

[0083] Regularly mine new synonyms from the user chat records (such as after the user asks "how to offset accounts" and then asks "reimbursement process");

[0084] Generate potential synonym candidates through a large language model and store them in the database after manual review.

[0085] Step 2: Perform synonym replacement on the user input question to generate an extended set of query statements;

[0086] For example, the extended set of query statements:

[0087] original_query (original question) = "How to write off out-of-town business trips";

[0088] expanded_queries (extended questions) = ["How to reimburse out-of-town business trips", "How to offset accounts for out-of-town business trips", "Settlement process for cross-provincial official business travel expenses"].

[0089] The actual policy document uses "Out-of-town business trip reimbursement process";

[0090] Step 3: Perform hybrid retrieval on each extended query statement respectively to generate the final candidate list.

[0091] Step 4: Replace the synonyms in the final candidate list with the main words;

[0092] For example, the actual policy document uses "Out-of-city Business Trip Reimbursement Process"; replace "Reimbursement Process" in the FAQ with "Reimbursement Process".

[0093] S5. Screen the Top-N candidate Q&A in the candidate list, and determine whether the Top-1 candidate Q&A is a personalized question. If so, call the structured database of the Finance Department to generate a query result based on the Top-1 candidate Q&A, and reply according to the query result and the Top-N candidate Q&A; otherwise, directly reply according to the Top-N candidate Q&A; where N is the screening threshold; where N is the screening threshold, generally set to 5.

[0094] Screen the Top-N candidate Q&A in the candidate list: Use the pre-trained re-ranking model to calculate the semantic matching degree between the candidate FAQ and the standardized query statement, sort the candidate FAQ in descending order according to the semantic matching degree, and screen the Top-N candidate Q&A after the descending order, so as to improve the semantic sorting effect.

[0095] Personalized questions are questions such as how much is the user's tuition fee, how much is the make-up fee to be paid, how much is the amount applied for reimbursement in the system, how much is the labor payment amount, etc., which need to be answered with specific numerical values.

[0096] Calling the structured database of the Finance Department to generate query results includes:

[0097] S51. Generate the query keywords for the Top-1 candidate Q&A.

[0098] The query keywords for generating the Top-1 candidate Q&A include:

[0099] S511. Call the large language model to perform intent recognition on the Top-1 candidate Q&A to obtain the user's intent.

[0100] S512. Call the large language model to extract the user's current input question and historical chat records according to the user's intent; through the query keywords, the specific financial matter data information of the user can be accurately matched (for example, keywords information such as expert funds matters, out-of-city reimbursement, credits, etc.).

[0101] S513. If the keywords required for the Top-1 candidate Q&A are extracted, output a prompt word that the data query keywords are complete: "The data query keywords are complete, no need for the user to supplement."; if the keywords required for the Top-1 candidate Q&A are not extracted, output a prompt word that the data query keywords are incomplete: "The data query keywords are incomplete, the user needs to supplement.", use the large language model (GLM-4-PLUS) to ask the user a rhetorical question according to the output prompt word and the user's intent, and return to step S2 to wait for the user to input.

[0102] For example, the top-1 candidate Q&A is: "How much can I reimburse?" In this question, the user does not specify what kind of reimbursement fee to query. Therefore, the large language model can be used to determine what money the user wants to consult based on the user's historical chat records. If tuition fees, book fees, accommodation fees, etc. were mentioned above, keywords can be directly extracted from this record. If it is impossible to determine what kind of reimbursement fee to query through intent recognition (it may be the first time to consult), then the system needs to ask a counter-question at this time.

[0103] S52. Call the large language model to generate a database query statement according to the query keywords;

[0104] S53. Check whether the database query statement generated by the large model is complete according to the fixed Python code. If so, execute step S54; otherwise, obtain an error message. If the number of times the database query statement is generated is greater than 3, directly set the query result as a fixed formula; otherwise, return to step S52 and call the large language model to regenerate the database query statement according to the error message and query keywords.

[0105] Python code is very flexible and different codes can be used for different scenarios. The Python pseudo-code example for checking the query statement output by the large model is as follows:

[0106] Input: The database query statement `query` generated by the large model

[0107] Setting: The predefined database table structure `TABLES`

[0108] - Each table contains field names and data types (for example, the `students` table contains `id`, `name`, `age` fields)

[0109] Step 1: **Check the basic structure of the query statement**

[0110] - If the query statement does not start with "SELECT" or does not contain a "FROM" clause, return an error:

[0111] - Error message: "The query statement must contain SELECT and FROM clauses"

[0112] Step 2: **Extract the query fields**

[0113] - Extract the field name part between "SELECT" and "FROM" from the query statement

[0114] - If the field part is empty, return an error:

[0115] - Error message: "No query fields are specified"

[0116] - If the field name is in the incorrect format (e.g., misspelled field name, duplicate fields, etc.), return an error:

[0117] - Error message: "The query field format is incorrect"

[0118] Step 3: **Extract the table name and verify the table structure**

[0119] - Extract the table name after the "FROM" clause in the query statement

[0120] - If the table name does not exist in the predefined table structure `TABLES`, return an error:

[0121] - Error message: "The table 'table_name' does not exist in the database"

[0122] Step 4: **Field validity check**

[0123] - For each field name extracted from the "SELECT" clause, check if it exists in the field list of the corresponding table

[0124] - Check each field:

[0125] - If the field exists and is spelled correctly, mark it as a valid field

[0126] - If the field is misspelled or does not exist in the table, return an error:

[0127] - Error message: "The field 'field_name' does not exist in the table 'table_name'"

[0128] - Record all invalid fields and provide specific error messages for each invalid field

[0129] Step 5: **Query condition validation (optional)**

[0130] - If the query statement contains a WHERE clause, check if the fields in the query conditions are valid in the table structure

[0131] - If the conditional field does not exist or is misspelled, return an error:

[0132] - Error message: "The field 'field_name' in the WHERE clause does not exist in the table 'table_name'"

[0133] Step 6: **Generate correction suggestions**

[0134] - Generate corresponding correction suggestions based on the error type:

[0135] - If it is a table name error, suggest checking the table name spelling or verifying the existence of the table

[0136] - If it is a field name error, it is recommended to verify the field name spelling or check whether the field exists in the table structure

[0137] - If it is a field error in the WHERE clause, it is recommended to check whether the field is correct or confirm the type of the query condition field

[0138] Output:

[0139] - If the query statement is valid, return "Query valid"

[0140] - If the query statement is invalid, return the error message and correction suggestions

[0141] Example:

[0142] - Input: Query statement `SELECT name,salary FROM teachers WHERE salary>5000`

[0143] - Output: Error message: "The field'salary' does not exist in the table'students'", correction suggestion: "Please verify the field name spelling or check whether the field exists in the table 'teachers'".

[0144] S54. Dynamically query the financial department database according to the database query statement to obtain the query result.

[0145] Repeatedly verify the identity information of students or teachers through the account information, ID number, student number, etc. of the user login until the verification passes or reaches the preset verification times. When the verification passes, dynamically query the financial department database according to the database query statement to obtain the query result (i.e., the numerical answer). When the verification fails and reaches the preset verification times, set the query result to a fixed phrase; the fixed phrase is "No relevant information was found. Please guide the user to operate according to the business knowledge for answering".

[0146] The solution reply includes: setting the Prompt reply template, and calling the large language model to generate a reply using the query result, Top-N candidate Q&A, and the preset Prompt template.

[0147] In the embodiment of the present invention, the used Prompt template is:

[0148] ## Role

[0149] Your identity is a consulting expert in the financial department of a university, and your name is Caizhizhi.

[0150] ## Task

[0151] Please think and analyze step by step, and answer the user's question based on the **user's question**, **user's intention**, **user's information** and **business knowledge**.

[0152] ## User Information

[0153] Fixed Identity Information

[0154] Identity: Student | Teacher

[0155] Name: XX

[0156] ID Number: XX

[0157] Student ID | Employee ID: XX

[0158] Dynamically Queryable Information

[0159] Credits: XX (compulsory), X (elective), and the college's minimum of 30 credits must be met.

[0160] Tuition Fees: Student XX has not paid tuition fees of XX yuan for the XX academic year.

[0161] Reimbursement Expenses: XX

[0162] Amount of Labor that can be Allocated: XX

[0163] Project Funds: XX

[0164] ## Time

[0165] Current Beijing Time: XX year XX month XX day XX:XX

[0166] ## Business Knowledge

[0167] 1. How much can I reimburse?

[0168] Answer: XX

[0169] 2. How to reimburse?

[0170] Answer: XX

[0171] 3. Who can be reimbursed?

[0172] Answer: XX

[0173] 4. What's the difference between business trips and official expenses?

[0174] Answer: XX

[0175] 5. What are the policies for in-city and out-of-city tourism?

[0176] Answer: XX

[0177] ## User Chat History

[0178] User: XX

[0179] Intelligent Assistant: XX

[0180] User: XX

[0181] ## User's Latest Intention

[0182] XX

[0183] On the other hand, the present invention adopts a financial intelligent question - answering system based on hybrid retrieval and dynamic query. This system is used to execute the above - mentioned financial intelligent question - answering method based on hybrid retrieval and dynamic query, and includes:

[0184] A multimodal intention recognition module, which is used to call a large - language model to recognize the intention of the question currently input by the user and the historical chat records;

[0185] A semantic rewriting module, which is used to call a large - language model to semantically rewrite the question input by the user in combination with the true intention of the user to generate a standardized query statement;

[0186] A hybrid retrieval and re - ranking module, which is used to obtain a candidate list from the FAQ knowledge base in a hybrid retrieval mode according to the standardized query statement and the user intention classification result; if there is no candidate Q&A in the candidate list, a standardized answer is returned; otherwise, the top - N candidate Q&As are screened from the candidate list;

[0187] A dynamic database query module: determines whether the top - 1 candidate Q&A is a personalized question. If so, it calls the structured database of the finance department to generate a query result;

[0188] A reply module: calls a large model to generate a reply by using the query result, the top - N candidate Q&As, and a pre - set Prompt template.

[0189] The above - mentioned embodiments further elaborate on the purpose, technical solutions, and advantages of the present invention. It should be understood that the above - mentioned embodiments are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made to the present invention within the spirit and principles of the present invention shall be included within the protection scope of the present invention.

Claims

1. A financial intelligent Q&A method based on hybrid retrieval and dynamic query, characterized in that Including: S1. Build an FAQ knowledge base and a structured database of the finance department; S2. Receive the question currently input by the user, retrieve the user's historical chat records, call the large language model to perform intent recognition on the question currently input by the user and the historical chat records, and obtain the user's true intent and the user intent classification result; S3. Call the large language model to semantically rewrite the question currently input by the user in combination with the user's true intent, and generate a standardized query statement; S4. Use a hybrid retrieval mode to obtain a candidate list from the FAQ knowledge base according to the standardized query statement and the user intent classification result; if there is no candidate Q&A in the candidate list, reply with a standardized answer; Otherwise, execute step S5; S5. Screen the Top-N candidate Q&As in the candidate list, and determine whether the Top-1 candidate Q&A is a personalized question. If so, call the structured database of the finance department to generate a query result according to the Top-1 candidate Q&A, and reply according to the query result and the Top-N candidate Q&As; otherwise, directly reply according to the Top-N candidate Q&As; where N is the screening threshold.

2. The financial intelligent question-answering method based on hybrid retrieval and dynamic query according to claim 1, wherein Performing intent recognition on the question currently input by the user and the historical chat records includes: determining whether the question currently input by the user contains a picture. If so, use a multimodal large model to parse the question currently input by the user to obtain structured semantic information; call the large language model and combine the structured semantic information and the user's historical chat records to identify the user's true intent and the user intent classification result; otherwise, directly call the large language model and combine the question currently input by the user and the user's historical chat records to identify the user's true intent and the user intent classification result.

3. A financial intelligent question-answering method based on hybrid retrieval and dynamic query according to claim 1, characterized in that, Obtaining a candidate list from the FAQ knowledge base includes: S41. Set a threshold Q, and calculate the semantic matching similarity between the standardized query statement and the Q&A in the FAQ knowledge base using semantic retrieval and keyword retrieval respectively and the keyword matching similarity where i is the index of the Q&A in the FAQ knowledge base; S42. Use the question and answer in the FAQ knowledge base whose keyword matching similarity with the standardized query statement is not less than the threshold Q as the candidate question and answer for keyword retrieval, and sort the candidate question and answer for keyword retrieval according to the keyword matching similarity to obtain the candidate list L2 for keyword retrieval; S43. Use the user intention classification result to perform weighted combination on the semantic matching similarity and the keyword matching similarity of the Q&A in the FAQ knowledge base to obtain the final matching similarity of the Q&A in the FAQ knowledge base S44. Use the Q&A in the FAQ knowledge base whose matching similarity with the standardized query statement is not less than the threshold Q as the candidate Q&A for hybrid retrieval, and sort the candidate Q&A for hybrid retrieval according to the matching similarity to obtain the candidate list L3 for hybrid retrieval; S45. Use the conflict detection rule to adjust the candidate list L3 of the hybrid retrieval according to the candidate list L2 retrieved by keywords, and obtain the final candidate list.

4. The financial intelligent Q&A method based on hybrid retrieval and dynamic query according to claim 3, characterized in that, The semantic matching similarity of the Q&A in the FAQ knowledge base and the keyword matching similarity are weighted and combined, including: Among them, α is the keyword retrieval weight, and β is the semantic retrieval weight. They are respectively the basic weights of keyword retrieval and semantic retrieval for each type of intention j, and p j represents the probability that the user intention belongs to the j-th type of intention, σ() is the Sigmoid function, and J is the number of intention categories.

5. The financial intelligent question-answering method based on hybrid retrieval and dynamic query according to claim 3, wherein Adjusting the candidate list of the hybrid retrieval by the candidate list retrieved by keywords includes: S451. Build a proprietary noun library in the financial field and a priority ranking table of proprietary nouns in the financial field; S452. Extract proprietary nouns from the candidate lists L2 and L3 of keyword retrieval and hybrid retrieval respectively according to the proprietary noun library in the financial field, and obtain the proprietary noun sets of the candidate lists L2 and L3; S453. Combine the proper nouns that belong to the proper noun set of candidate list L2 but do not belong to the proper noun set of candidate list L3 to obtain a set of conflicting proper nouns E = {e1, e2, …, e M}, and select the proper noun with the highest priority in the proper noun set E according to the proper noun priority sorting table in the financial field Among all candidate Q&As in candidate list L2 that contain the proper noun select the candidate Q&A with the highest keyword matching similarity and insert it into candidate list L3 to obtain the final candidate list; where e m is a proper noun in the proper noun set E, and M is the number of proper nouns in the proper noun set E.

6. A financial intelligent question answering method based on hybrid retrieval and dynamic query according to claim 1, characterized in that, Calling the structured database of the finance department to generate a query result includes: S51. Generate the query keywords for the Top-1 candidate Q&A; S52. Call the large language model to generate a database query statement according to the query keywords; S53. According to the fixed Python code, detect whether the database query statement generated by the large model is complete. If so, execute step S54; otherwise, obtain an error message. If the number of times of generating the database query statement is greater than 3, directly set the query result as a fixed formula; otherwise, return to step S52, and call the large language model to regenerate the database query statement according to the error message and the query keywords; S54. Dynamically query the finance department database according to the database query statement to obtain a query result.

7. A financial intelligent Q&A method based on hybrid retrieval and dynamic query according to claim 6, characterized in that Generating the query keywords for the Top-1 candidate Q&A includes: S511. Invoke the large language model to perform intent recognition on the top-1 candidate Q&A to obtain the user's intent; S512. Invoke the large language model to extract keywords from the question currently input by the user and the historical chat records according to the user's intent; S513. If the keywords required for the top-1 candidate Q&A are extracted, output a prompt indicating that the data query keywords are complete; if the keywords required for the top-1 candidate Q&A are not extracted, output a prompt indicating that the data query keywords are incomplete; use the large language model to ask the user a rhetorical question according to the output prompt and the user's intent, and return to step S2 to wait for the user to input.

8. A financial intelligent question answering method based on hybrid retrieval and dynamic query according to claim 6, characterized in that, Dynamically query the finance department database according to the database query statement, including: repeatedly verifying the user's identity information through the information logged in by the user until the verification passes or the preset verification times are reached; when the verification passes, dynamically query the finance department database according to the database query statement to obtain the query result; when the verification fails and the preset verification times are reached, set the query result to a fixed formula.

9. A financial intelligent question-answering method based on hybrid retrieval and dynamic query according to claim 1, characterized in that, Reply according to the top-N candidate Q&As, including: setting a Prompt reply template, and invoking the large language model to generate a reply according to the top-N candidate Q&As and the Prompt template.

10. A financial intelligent question-answering system based on hybrid retrieval and dynamic query, which is used to execute a financial intelligent question-answering method based on hybrid retrieval and dynamic query as described in any one of claims 1 to 9, characterized in that, Including: A multimodal intent recognition module for invoking the large language model to perform intent recognition on the question currently input by the user and the historical chat records; A semantic rewriting module for invoking the large language model to semantically rewrite the question input by the user in combination with the user's true intent to generate a standardized query statement; A hybrid retrieval and reordering module for obtaining a candidate list from the FAQ knowledge base in a hybrid retrieval mode according to the standardized query statement and the user intent classification result; if there is no candidate Q&A in the candidate list, reply with a standardized answer; Otherwise, screen the top-N candidate Q&As from the candidate list; A dynamic database query module: determine whether the top-1 candidate Q&A is a personalized question, if so, invoke the finance department structured database to generate a query result; A reply module: invoke the large model to generate a reply using the query result, the top-N candidate Q&As, and the preset Prompt template.

Citation Information

Cited By

  • Retrieval question and answer method and device for table, medium, equipment and program product

    CN120448407A

  • Retrieval question and answer method, device, medium, equipment and program product for table

    CN120448407B

  • Data analysis question and answer platform based on large model and knowledge vector library

    CN120448508A

  • Method for service recommendation based on identified user intention, computing device and storage medium

    CN120632219A

  • Intelligent customer service processing method, device and equipment based on large language model

    CN120892544A