Question answering method and device based on retrieval enhancement generation, equipment and medium

By enhancing the question-answering method of the large language model through keyword and text vector retrieval in the knowledge base of the financial field, the shortcomings of traditional models in understanding professional terminology and timeliness are solved, providing more accurate and timely access to financial information.

CN120950640APending Publication Date: 2025-11-14AGRICULTURAL BANK OF CHINA
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202511046945.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-29
Publication Date
2025-11-14

AI Technical Summary

Technical Problem

Traditional large language models cannot understand technical terms in financial question answering, resulting in inaccurate and lagging output results.

Method used

By using a retrieval-enhanced generation method, keyword and text vector retrieval is performed using a pre-established knowledge base, and a large language model is combined to process the target question, resulting in more accurate and timely answers.

Benefits of technology

This has enabled the output of large language models to be more timely and accurate, improving the professionalism and timeliness of information acquisition in the financial field.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120950640A_ABST
    Figure CN120950640A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a question answering method and device based on retrieval enhancement generation, equipment and a medium, and relates to the technical field of large language model question answering. The method comprises the steps of obtaining a target question input by a user; on the basis of keywords in the target question, keyword retrieval is carried out on each text in a pre-established knowledge base; each text in the knowledge base corresponds to a text vector; retrieving a text vector similar to the target vector in a pre-established knowledge base, and determining a text corresponding to the retrieved text vector; the target vector is a vector generated based on a target problem; and processing the target question and a retrieval result obtained in the knowledge base based on the large language model to obtain an answer corresponding to the target question. According to the technical scheme, text retrieval and vector retrieval are carried out in the knowledge base, the retrieval result with higher timeliness and higher accuracy is obtained, and then the answer output by the large language model can be obtained based on the retrieval result and the target question.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of large-scale language question answering technology, and in particular to a question answering method, apparatus, device and medium based on retrieval enhancement generation. Background Technology

[0002] With the accelerating digital transformation of the financial industry, massive amounts of financial data are growing exponentially. Investors and financial professionals face the daily challenge of accessing and understanding multi-dimensional information, including financial product information, market dynamics, and regulatory information.

[0003] When users want to acquire and understand multi-dimensional information in the financial field, they can ask questions to the large language model to obtain the financial information output by the large language model. Traditional large language models often cannot understand the meaning of professional financial terms, resulting in inaccurate output results; furthermore, because large language models rely on training data when generating answers, the output answers are lagging. Summary of the Invention

[0004] This invention provides a question-answering method, apparatus, device, and medium based on retrieval-enhanced generation, which can make the output of large language models more timely and accurate based on retrieval-enhanced generation.

[0005] According to one aspect of the present invention, a question-answering method based on retrieval enhancement generation is provided, the method comprising:

[0006] The target problem for obtaining user input;

[0007] Based on the keywords in the target question, keyword retrieval is performed on each text in a pre-established knowledge base; the knowledge base stores financial knowledge texts and internal investment banking texts; each text corresponds to a text vector;

[0008] Retrieve text vectors similar to the target vector from a pre-established knowledge base, and determine the text corresponding to the retrieved text vectors; the target vector is a vector generated based on the target question.

[0009] The target question and the retrieval results obtained from the knowledge base are processed based on a large language model to obtain the answer to the target question.

[0010] According to another aspect of the present invention, a question-answering apparatus based on retrieval enhancement generation is provided, comprising:

[0011] The target question acquisition module is used to acquire the target question input by the user.

[0012] The keyword retrieval module is used to perform keyword retrieval on each text in a pre-established knowledge base based on keywords in the target question; the knowledge base stores financial knowledge texts and internal investment banking texts; each text corresponds to a text vector;

[0013] The vector retrieval module is used to retrieve text vectors similar to the target vector from a pre-established knowledge base and determine the text corresponding to the retrieved text vectors; the target vector is a vector generated based on the target question.

[0014] The answer generation module is used to process the target question and the retrieval results obtained from the knowledge base based on the large language model to obtain the answer corresponding to the target question.

[0015] According to another aspect of the present invention, an electronic device is provided, the electronic device comprising:

[0016] At least one processor; and

[0017] A memory communicatively connected to the at least one processor; wherein,

[0018] The memory stores a computer program that can be executed by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform the question-answering method based on retrieval enhancement generation as described in any embodiment of the present invention.

[0019] According to another aspect of the present invention, a computer-readable storage medium is provided, the computer-readable storage medium storing computer instructions for causing a processor to execute and implement the question-answering method based on retrieval enhancement generation as described in any embodiment of the present invention.

[0020] The technical solution of this application includes: acquiring a target question input by a user; performing keyword retrieval on each text in a pre-established knowledge base based on keywords in the target question; the knowledge base stores financial knowledge texts and internal investment banking texts; each text corresponds to a text vector; retrieving text vectors similar to the target vector in the pre-established knowledge base, and determining the text corresponding to the retrieved text vector; the target vector is a vector generated based on the target question; and processing the target question and the retrieval results obtained in the knowledge base based on a large language model to obtain the answer corresponding to the target question. This technical solution, based on retrieval enhancement generation technology, performs text retrieval and vector retrieval in the knowledge base, obtaining more timely and accurate retrieval results, and then, based on the retrieval results and the target question, obtaining a more accurate answer output by the large language model.

[0021] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of the present invention, nor is it intended to limit the scope of the invention. Other features of the invention will become readily apparent from the following description. Attached Figure Description

[0022] To more clearly illustrate the technical solutions in the embodiments of this application, the drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0023] Figure 1 This is a flowchart of a question-answering method based on retrieval enhancement provided in Embodiment 1 of this application;

[0024] Figure 2 This is a flowchart of a question-answering method based on retrieval enhancement provided in Embodiment 2 of this application;

[0025] Figure 3 This is a schematic diagram of the architecture of a question-answering system based on retrieval enhancement generation according to Embodiment 2 of this application;

[0026] Figure 4 This is a schematic diagram of a question-answering device based on retrieval enhancement generation according to Embodiment 3 of this application;

[0027] Figure 5 This is a schematic diagram of the structure of an electronic device that implements a retrieval-enhanced question-answering method according to an embodiment of this application. Detailed Implementation

[0028] To enable those skilled in the art to better understand the present invention, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.

[0029] It should be noted that the terms "first," "second," "target," etc., used in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0030] Example 1

[0031] Figure 1 This application provides a flowchart of a search-enhanced question-answering method according to Embodiment 1. This embodiment is applicable to generating answers to user-input questions. The method can be executed by a search-enhanced question-answering device, which can be implemented in hardware and / or software and can be configured in an electronic device with data processing capabilities. Figure 1 As shown, the method includes:

[0032] S110, Obtain the target question input by the user;

[0033] The target question is what a user should input on the relevant query interface when they need to look up knowledge in the financial field or internal documents of an investment bank; the user can be an employee of the investment bank.

[0034] S120: Based on the keywords in the target question, perform keyword retrieval on each text in a pre-established knowledge base.

[0035] The knowledge base contains financial knowledge texts and internal investment banking texts; each text corresponds to a text vector; the knowledge base includes, but is not limited to, content related to: investment banking business knowledge, financial information, research reports, regulatory policies, internal systems, and internal regulations.

[0036] Specifically, after obtaining the target question, keywords are extracted from it, and the extracted keywords are used to perform keyword retrieval on the text in the knowledge base to obtain the text that matches the keywords.

[0037] S130, retrieve text vectors similar to the target vector from a pre-established knowledge base, and determine the text corresponding to the retrieved text vectors; the target vector is a vector generated based on the target question.

[0038] Specifically, after obtaining the target question, a vector transformation is performed on the target question to obtain the target vector. Since the target vector reflects the semantic information of the target question, text vectors similar to the target vector are retrieved from the text vectors in the knowledge base to obtain the retrieved text vector. In the knowledge base, since there is a one-to-one correspondence between text vectors and texts, the text corresponding to the retrieved text vector can be determined.

[0039] S140, Based on the large language model, the target question and the retrieval results obtained in the knowledge base are processed to obtain the answer corresponding to the target question.

[0040] Specifically, in steps S120 and S130, texts related to the target question are retrieved from the knowledge base. Then, among these texts, texts with a strong correlation to the target question are identified. These texts are combined with the target question to form new prompt words. Then, the prompt words are input into the large language model to obtain the answer to the target question output by the large language model.

[0041] The technical solution of this application includes: acquiring a target question input by a user; performing keyword retrieval on each text in a pre-established knowledge base based on keywords in the target question; the knowledge base stores financial knowledge texts and internal investment banking texts; each text corresponds to a text vector; retrieving text vectors similar to the target vector in the pre-established knowledge base, and determining the text corresponding to the retrieved text vector; the target vector is a vector generated based on the target question; and processing the target question and the retrieval results obtained in the knowledge base based on a large language model to obtain the answer corresponding to the target question. This technical solution, based on retrieval enhancement generation technology, performs text retrieval and vector retrieval in the knowledge base, obtaining more timely and accurate retrieval results, and then, based on the retrieval results and the target question, obtaining a more accurate answer output by the large language model.

[0042] Example 2

[0043] Figure 2 This is a flowchart of a question-answering method based on retrieval enhancement provided in Embodiment 2 of this application. This embodiment is an optimization based on the above embodiment.

[0044] like Figure 2 As shown, the method in this embodiment of the application specifically includes the following steps:

[0045] S210, Obtain the target question input by the user.

[0046] S220, based on a pre-established financial security rule base, perform security verification on the target question; if the security verification result meets the security requirements, then based on the keywords in the target question, perform keyword retrieval on each text in a pre-established knowledge base.

[0047] The knowledge base contains financial knowledge texts and internal investment banking texts; each text corresponds to a text vector.

[0048] For example, the target question input by the user is subject to security verification. The security verification is based on a financial security rule base, which is set up based on relevant regulations in the financial field and internal company regulations. This financial security rule base is continuously updated.

[0049] In a specific example, when performing security verification on the target question based on a pre-established financial security rule base, if the input content contains violent, pornographic, or politically sensitive content, a security prompt will be issued: "The content you entered may violate security regulations. Please re-enter." Then, the target question entered by the user can be retrieved again until the security verification passes. Based on the keywords in the target question, keyword retrieval will be performed on each text in the pre-established knowledge base.

[0050] This solution is configured in this way to prevent non-compliant content from appearing in user questions and to ensure the health of the query environment.

[0051] Optionally, in this embodiment of the application, the knowledge base update process includes: continuously acquiring a first type of document and converting the first type of document into a fourth text; the first type of document is a document that meets the high update frequency requirement; processing the fourth text based on a hash algorithm to obtain a unique text identifier; if there is no text in the knowledge base that has the same document identifier as the first type of document, then inserting the fourth text and the fourth text vector into the knowledge base; if there is text in the knowledge base that has the same document identifier as the first type of document, but the unique text identifier of the text is different from the unique text identifier of the fourth text, then replacing the corresponding text and text vector in the knowledge base with the fourth text and the fourth text vector.

[0052] This application categorizes documents requiring updates into a first category and a second category. The first category refers to documents with a high update frequency, while the second category refers to documents with a low update frequency. For example, the first category could be financial information or similar data that needs to be updated as frequently as possible. The second category consists of internal company documents that can be updated periodically.

[0053] For the first type of document, first-type documents are continuously acquired. If a first-type document is received, it is converted into a fourth type of text. This fourth type of text is then processed using a hash algorithm to obtain a unique text identifier. It should be noted that the document identifier is not unique; for example, two documents with the same document identifier may have different content. Similarly, each text in the knowledge base has a corresponding document identifier and a unique text identifier. The document identifiers in the knowledge base can be searched. If no document identifier for a received first-type document is found, the fourth text and its vector are inserted into the knowledge base according to the category. If a text in the knowledge base has the same document identifier as a received first-type document, its unique text identifier is compared with that of the first-type document. If they do not match, their contents are different. An index can be established between the vector-converted fourth text vector and the fourth text, and the corresponding text vector and text in the knowledge base are replaced with the fourth text vector and its text.

[0054] It should be noted that when converting documents (including first-class and second-class documents) into text, since the document formats may not be uniform, they can be converted into text first, and then long texts can be divided into text segments of appropriate length based on logical structures such as paragraphs and chapters and preset character length thresholds. Then, a text vector is generated for each text segment, and an index is established between the text vector and the corresponding text before storing it in the knowledge base.

[0055] This solution is configured in such a way that the first type of documents can be continuously updated, and duplicate documents can be avoided from being added to the knowledge base again during the update process.

[0056] Optionally, in this embodiment of the application, the knowledge base update process includes: if a second type of document is received, obtaining the version of the second type of document and the version of the corresponding text in the knowledge base; the second type of document is a document that meets the low update frequency requirement; if the version of the second type of document is an updated version, converting the second type of document into text, and saving the text and the corresponding text vector to the knowledge base.

[0057] Specifically, for the second type of document, the version of the second type of document can be obtained, and then it can be determined whether the version of the second type of document is consistent with the version of the corresponding text in the knowledge base. If they are consistent, there is no need to update the knowledge base. If they are inconsistent and the version of the second type of document is updated, the second type of document is converted into text, and the text and the corresponding text vector are saved to the knowledge base.

[0058] This solution allows for updating documents with version numbers in the knowledge base that have a low update frequency, thus saving corresponding computing resources.

[0059] S230, retrieve text vectors similar to the target vector from a pre-established knowledge base, and determine the text corresponding to the retrieved text vectors; the target vector is a vector generated based on the target question.

[0060] S240, Based on the keyword retrieval score of the first text and the vector similarity score between the first text vector and the target vector, determine the second text in the first text; the first text is the text retrieved from the knowledge base.

[0061] The first text is the text obtained by keyword retrieval in the knowledge base and the text corresponding to similar text vectors obtained by retrieval based on the target vector.

[0062] The keyword retrieval score for the first text is based on a keyword matching (e.g., the BM25 algorithm) scoring mechanism using term frequency and inverse document frequency, primarily used to determine the degree of match between the document and the query. It assesses relevance by calculating the frequency of the query term in the document and the rarity of that term in the entire corpus. The more frequently the query term appears in the document, and the rarer the term, the more relevant the document is.

[0063] The vector similarity score between the first text vector and the target vector can be calculated based on the similarity between the first text vector and the target vector. This similarity can be cosine similarity, etc. The higher the vector similarity between the first text vector and the target vector, the higher the vector similarity score between the first text vector and the target vector.

[0064] Furthermore, after determining the keyword retrieval score of the first text and the vector similarity score between the first text vector and the target vector, a second text that is more closely matched to the target question is determined from the first text based on the scores. This second text is the information retrieved in the retrieval enhancement generation technique. This scheme is designed in such a way that the determined second text has a higher degree of matching with the target question, resulting in higher quality answers generated subsequently.

[0065] In this embodiment of the application, optionally, determining a second text in the first text based on the keyword retrieval score of the first text and the vector similarity score between the first text vector and the target vector includes: performing a weighted summation of the keyword retrieval score and the vector similarity score to obtain the total retrieval score corresponding to the first text; sorting the first text according to the total retrieval score; determining a third text in the first text based on the sorting result; and determining a second text in the third text based on the degree of matching between the third text and the target question.

[0066] Specifically, the keyword retrieval score is set as the first weight, and the vector similarity score is set as the second weight. The sum of the first weight and the second weight is 1. The result of multiplying the first weight by the keyword retrieval score and the result of multiplying the second weight by the vector similarity score are superimposed to obtain the total retrieval score. The first text is sorted according to the total retrieval score, and a third text is determined from the first text according to the sorting result. The second text is determined from the third text according to the degree of matching between the third text and the target question.

[0067] In this embodiment of the application, optionally, determining the second text from the third text based on the matching degree between the third text and the target question includes: processing the concatenated sequence based on a pre-trained cross-encoder to obtain a similarity score between the target question and each third text; the concatenated sequence is obtained by concatenating the target question and the third text; sorting the third text according to the similarity score, and determining the second text from each third text according to the sorting result.

[0068] This scheme is designed to perform two rounds of filtering on the first text. First, based on the total search score, a third text with a higher total score is selected from the first text. Then, based on a cross-encoder, a second text with a higher degree of matching with the target question is selected from the third text. This setup makes the second text more relevant to the target question, thereby making the subsequently generated answer more accurate.

[0069] S250, combine the second text with the target question to obtain the target prompt word.

[0070] S260, The target prompt words are processed based on the large language model to obtain the answer corresponding to the target question.

[0071] S270 processes the answer to the target question based on a pre-trained compliance check model. If the output result shows that the answer is compliant, the answer is processed for sensitive information so that the processed answer can be fed back to the user.

[0072] Specifically, after generating the answer to the target question, a text classification model is trained based on a compliance knowledge base to automatically classify the generated answer. If the answer is non-compliant, a friendly prompt is given: "The generated content does not meet relevant regulatory requirements; please regenerate." If the generated content is compliant, sensitive content processing is required, using a sensitive content rule base to identify, filter, and replace sensitive content. Furthermore, the generated answer includes a source link. During the answer generation process, the source of each information fragment is recorded. When needed, users can click the source link to view the original document information upon which the specific content of the answer is based, facilitating further verification and in-depth understanding of the relevant knowledge.

[0073] The technical solution of this application focuses on the storage and management of core knowledge in the investment banking field. The knowledge base can collect the latest market dynamics, industry research, financial news, and other information in real time. It also includes various financial regulatory systems and regulations, providing a comprehensive and accurate knowledge source for intelligent question answering. Simultaneously, the knowledge base supports dynamic incremental updates, ensuring it remains highly up-to-date. The real-time update mechanism for the first type of document allows newly generated important information to be quickly integrated into the knowledge base. For example, major events in the financial market or newly issued regulations by regulatory authorities can be updated and used to answer user questions in a short time, greatly improving the timeliness of information. The batch, timed updates for the second type of document ensure the integrity of the knowledge base content. New knowledge content is regularly added over time, preventing the knowledge system from becoming outdated or incomplete.

[0074] This technical solution avoids sensitive and security issues at the input level by performing security verification on the target question. It also significantly improves the recall rate of technical terms by combining keyword matching (BM25) and semantic vector retrieval. Traditional single retrieval methods struggle to simultaneously achieve accurate keyword matching and semantic understanding, while the hybrid retrieval strategy in this application's embodiments achieves a synergistic effect between the two. When faced with complex technical terms in the investment banking field, it can retrieve relevant information from the knowledge base more comprehensively and accurately, providing users with more complete answers. For example, in questions involving complex financial derivatives, this retrieval strategy can avoid information omissions caused by semantic misunderstanding biases or limitations in keyword matching, helping users quickly obtain comprehensive and accurate knowledge.

[0075] This technical solution employs a newly added compliance detection layer to perform security verification and compliance checks on generated content, and supports source link queries—a design of significant importance. In the stringent regulatory environment of the financial industry, compliance is paramount. The compliance detection layer, based on a text classification model trained on a compliance and regulatory knowledge base, checks the generated content and automatically determines its compliance status. This effectively prevents intelligent question-and-answer systems from outputting content that violates financial regulations, thus preventing potential legal risks and reputational damage. Simultaneously, the source link query function provides users with assurance of information credibility. Users can view the original documents upon which the answers are based through source links, enhancing their trust in the answers. This also facilitates internal audits and knowledge management for financial institutions, further standardizing the use and dissemination of information.

[0076] In summary, the embodiments of this application innovate in many aspects, such as knowledge base construction and updating, model framework optimization, and content review, comprehensively improving the performance and reliability of intelligent question-answering systems in the investment banking field, and meeting the urgent needs of the financial industry for professional, efficient, and compliant intelligent question-answering services.

[0077] In a specific instance, refer to Figure 3 A schematic diagram of the architecture of a question-answering system based on retrieval enhancement is shown in the figure. The method described in this application embodiment is implemented through the following five stages:

[0078] The first stage is to establish a document library.

[0079] Collect documents related to investment banking business knowledge, financial information, research reports, and internal systems, categorize the documents, and form a professional document library.

[0080] For dynamic updates to the document library: For hot data (Category 1 documents), a document-level dynamic incremental real-time update method is adopted. A unique ID is generated for each document based on a hash algorithm, accurately identifying the document's content characteristics. During each knowledge indexing process, the system compares the unique ID of a new document with the IDs of existing documents. If it's a new ID, the document is considered new content, and the corresponding insertion operation is performed in the corresponding knowledge base according to its category. If the unique ID already exists but the document content has been updated, an update operation is performed. If the unique IDs are completely identical, the document is skipped to avoid duplicate storage.

[0081] Timed Batch Updates of Lesser-Known Facts (Category 2 Documents): For lesser-known facts, timed batch updates are performed based on different text categories. During updates, the latest lesser-known facts documents are retrieved from the data source, and the data collection, text segmentation, vectorization embedding, index building, and data import processes described above are repeated. New lesser-known facts content is added to the knowledge base in batches, ensuring the completeness and timeliness of the knowledge base.

[0082] The second stage involves building a knowledge base.

[0083] File preprocessing: Converting files of various formats into document objects.

[0084] Text segmentation: Utilizing professional text processing tools for segmentation. Based on logical structures such as paragraphs and chapters, and preset character length thresholds, long texts are divided into appropriately sized segments. For example, a financial research report spanning dozens of pages can be segmented into segments of 500-1000 words each, facilitating subsequent vectorization processing.

[0085] Vectorized embedding: The BERT model is used to vectorize the segmented text fragments, capturing the semantic information of words in the text and transforming them into vector representations in a high-dimensional space. This enables computers to understand the semantic features of the text, providing a foundation for subsequent index construction and retrieval.

[0086] Index building and data storage: Based on the vectorized text vectors, an efficient indexing algorithm, such as Faiss (Facebook AISimilarity Search), is used to build the index. Faiss can quickly calculate the similarity between vectors, greatly improving retrieval efficiency. The indexed text vectors and their corresponding original text fragments are then stored in the database.

[0087] Phase 3: Search.

[0088] When a user enters a question, a security check is first performed based on a security rule base to indicate the legitimacy and reasonableness of the input. Questions that pass the security check are then vectorized into text, and relevant content is retrieved from the knowledge base. The retrieval employs a hybrid approach combining keyword matching (BM25) and semantic vectors.

[0089] Keyword matching (BM25) combined with semantic vector retrieval is a retrieval strategy that integrates traditional information retrieval techniques with modern semantic retrieval techniques. Based on a scoring mechanism using term frequency and inverse document frequency, it is primarily used to determine the degree of match between a document and a query. It assesses relevance by calculating the frequency of the query term in the document and the rarity of that term in the entire corpus. In short, the more frequently a query term appears in a document, and the rarer that term, the more relevant the document.

[0090] Semantic vector retrieval: The query and the document are encoded as high-dimensional vectors, and semantic relevance is determined by calculating the similarity between the vectors (such as cosine similarity).

[0091] On one hand, the BM25 algorithm is used to calculate the keyword matching score between the question and documents in the knowledge base. The BM25 algorithm calculates the score based on factors such as the frequency of keyword occurrence in the documents and document length, identifying a set of documents with a high degree of matching with the question's keywords. On the other hand, the question is vectorized using the same vectorization model as when the knowledge base was constructed, resulting in a vector representation of the question. Next, the similarity (e.g., cosine similarity) between the question vector and the document vectors in the knowledge base is calculated to obtain a score based on semantic vector retrieval. Finally, the two scores are weighted and summed according to different weights, and the retrieval results are ranked according to the comprehensive score, returning the relevant document information with higher scores to the subsequent enhancement stage.

[0092] A parameter α is introduced as a weighting factor to control the weight allocation ratio of the two methods.

[0093] α = 1: Search is performed entirely based on semantic vectors, ignoring keyword vectors.

[0094] α = 0: Search is performed entirely based on keyword vectors, ignoring semantic vectors.

[0095] 0<α<1: Use a combination of keyword vectors and semantic vectors, with specific weights adjusted according to actual needs.

[0096] In mixed search, the mixed score can be calculated using the following formula:

[0097] hybrid_score=(1-α)*keyword_score+α*semantic_Score

[0098] Here, keyword_score is the score of the keyword vector, semantic_score is the score of the semantic vector, and hybrid_score is the weighted score of the two. In this invention, α = 0.5, that is, keyword search and semantic vector search each account for 50% of the weight.

[0099] Cross-Encoder Re-Ranking. A cross-encoder is a deep neural network that processes a query and a document as a pair of inputs to generate a score representing their relevance. Its main steps include:

[0100] Encoding and Scoring: The query and each retrieved document are concatenated together, and their relevance scores are calculated using a cross-encoder model.

[0101] Reordering: Based on the scores generated by the cross-encoder, the retrieved documents are reordered, with the most relevant documents placed first.

[0102] The advantage of cross-encoders lies in their ability to capture subtle semantic relationships between queries and documents, thereby generating more accurate relevance scores. However, due to their high computational cost, they are generally not suitable for direct retrieval of large-scale datasets, but are instead used for fine-tuning of preliminary search results.

[0103] By rapidly retrieving candidate documents through hybrid retrieval and then reordering them using a cross-encoder, the performance of the RAG system can be significantly improved. This combined approach leverages both the efficiency of hybrid retrieval and the accuracy of cross-encoders. In question-answering systems, hybrid retrieval can first extract question-related documents from the knowledge base, and then cross-encoders can be used to reorder these documents, ensuring that the most relevant documents are prioritized for answer generation. This approach further enhances the quality and relevance of the generated results while maintaining retrieval efficiency.

[0104] Fourth stage: Generating answers.

[0105] The question and the content of the top k relevant documents retrieved are combined to form a new prompt, which is then submitted to the large model LLM.

[0106] LLM (Language Modeling) generates the final text answer or content based on the input query and retrieved relevant information, utilizing its language generation capabilities. During generation, the model considers the semantics and context of the input information to produce a natural, fluent, and targeted response. After answer generation, a compliance check layer performs compliance checks on the generated content. A text classification model trained on a compliance knowledge base automatically categorizes the generated answers. If an answer is non-compliant, a friendly prompt is provided: "The generated content does not meet relevant regulatory requirements; please regenerate." If the generated content is compliant, sensitive content processing is required, using a sensitive rule base to identify, filter, and replace sensitive content. Furthermore, the system supports source link queries. During answer generation, the source of each information fragment is recorded. When needed, users can click on the source link to view the original document information upon which the specific content of the answer is based, facilitating further verification and in-depth understanding of relevant knowledge.

[0107] Phase 5: Application. The retrieval-enhanced question-and-answer method proposed in this invention is applied to inquiries regarding professional knowledge questions and answers, professional research, operational guidelines, market analysis, risk assessment, and compliance regulations for various business lines in investment banking, providing support for the standardization and efficiency of business operations.

[0108] Example 3

[0109] Figure 4 This is a schematic diagram of a question-answering device based on retrieval enhancement generation provided in Embodiment 3 of this application. This device can execute the question-answering method based on retrieval enhancement generation provided in any embodiment of this invention, and possesses the corresponding functional modules and beneficial effects of the method. For example... Figure 4 As shown, the device includes:

[0110] The target question acquisition module 310 is used to acquire the target question input by the user.

[0111] The keyword retrieval module 320 is used to perform keyword retrieval on each text in a pre-established knowledge base based on keywords in the target question; the knowledge base stores financial knowledge texts and internal investment banking texts; each text corresponds to a text vector;

[0112] The vector retrieval module 330 is used to retrieve text vectors similar to the target vector from a pre-established knowledge base, and to determine the text corresponding to the retrieved text vectors; the target vector is a vector generated based on the target question.

[0113] The answer generation module 340 is used to process the target question and the retrieval results obtained in the knowledge base based on the large language model to obtain the answer corresponding to the target question.

[0114] The technical solution of this application embodiment includes: a target question acquisition module 310, used to acquire a target question input by a user; a keyword retrieval module 320, used to perform keyword retrieval on each text in a pre-established knowledge base based on keywords in the target question; the knowledge base stores financial knowledge text and investment banking internal text; each text corresponds to a text vector; a vector retrieval module 330, used to retrieve text vectors similar to the target vector in the pre-established knowledge base, and determine the text corresponding to the retrieved text vector; the target vector is a vector generated based on the target question; and an answer generation module 340, used to process the target question and the retrieval results obtained in the knowledge base based on a large language model to obtain the answer corresponding to the target question. This technical solution is based on retrieval enhancement generation technology, performing text retrieval and vector retrieval in the knowledge base to obtain more accurate retrieval results, and then, based on the retrieval results and the target question, a more accurate answer output by the large language model can be obtained.

[0115] Optionally, in this embodiment of the application, the answer generation module 340 includes:

[0116] The second text determination unit is used to determine the second text in the first text based on the keyword retrieval score of the first text and the vector similarity score between the first text vector and the target vector; the first text is the text retrieved from the knowledge base;

[0117] The target prompt word determination unit is used to combine the second text with the target question to obtain the target prompt word;

[0118] The answer generation unit is used to process the target prompt words based on a large language model to obtain the answer corresponding to the target question.

[0119] Optionally, in this embodiment of the application, the second text determination unit includes:

[0120] The total retrieval score determination subunit is used to perform a weighted summation of the keyword retrieval score and the vector similarity score to obtain the total retrieval score corresponding to the first text;

[0121] The third text determination subunit is used to sort the first text according to the total search score, and determine the third text in the first text according to the sorting result;

[0122] The second text determination subunit is used to determine the second text from the third text based on the degree of matching between the third text and the target question.

[0123] Optionally, in this embodiment of the application, the second text determining subunit is specifically used for:

[0124] The concatenated sequence is processed using a pre-trained cross-encoder to obtain similarity scores between the target question and each third text; the concatenated sequence is obtained by concatenating the target question and the third text.

[0125] The third text is sorted according to the similarity score, and the second text is determined from each third text based on the sorting result.

[0126] Optionally, in this embodiment of the application, the apparatus further includes: a first knowledge base update module, comprising:

[0127] The first type of document acquisition unit is used to continuously acquire first type of documents and convert the first type of documents into fourth text; the first type of documents are documents that meet the requirements of high update frequency.

[0128] A unique text identifier determination unit is used to process the fourth text based on a hash algorithm to obtain a unique text identifier;

[0129] The text storage unit is used to insert the fourth text and the fourth text vector into the knowledge base if there is no text in the knowledge base that has the same document identifier as the first type of document; if there is text in the knowledge base that has the same document identifier as the first type of document, but the unique text identifier of the text is different from the unique text identifier of the fourth text, the corresponding text and text vector in the knowledge base are replaced with the fourth text and the fourth text vector.

[0130] Optionally, in this embodiment of the application, the apparatus further includes: a second knowledge base update module, comprising:

[0131] The version determination unit is used to obtain the version of the second type of document and the version of the corresponding text in the knowledge base if a second type of document is received; the second type of document is a document that meets the low update frequency requirement.

[0132] The text storage unit is used to convert the second type of document into text if the version of the second type of document is an updated version, and to save the text and the corresponding text vector to the knowledge base.

[0133] Optionally, in this embodiment of the application, the device further includes:

[0134] The security verification module is used to perform security verification on the target question based on a pre-established financial security rule base; if the security verification result meets the security requirements, the module performs a keyword retrieval of each text in the pre-established knowledge base based on the keywords in the target question.

[0135] The compliance check module is used to process the answers to the target questions based on a pre-trained compliance check model. If the output result shows that the answer is compliant, the answer will undergo sensitive information processing, and the processed answer will be fed back to the user.

[0136] The question-answering device based on retrieval enhancement provided in this application embodiment can execute the question-answering method based on retrieval enhancement provided in any embodiment of the present invention, and has the corresponding functional modules and beneficial effects of the execution method.

[0137] Example 4

[0138] Figure 5 A schematic diagram of an electronic device 10 that can be used to implement embodiments of the present invention is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices (e.g., helmets, glasses, watches, etc.), and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the invention described and / or claimed herein.

[0139] like Figure 5 As shown, the electronic device 10 includes at least one processor 11 and a memory, such as a read-only memory (ROM) 12 or a random access memory (RAM) 13, communicatively connected to the at least one processor 11. The memory stores computer programs executable by the at least one processor. The processor 11 can perform various appropriate actions and processes based on the computer program stored in the ROM 12 or loaded from storage unit 18 into the RAM 13. The RAM 13 may also store various programs and data required for the operation of the electronic device 10. The processor 11, ROM 12, and RAM 13 are interconnected via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.

[0140] Multiple components in electronic device 10 are connected to I / O interface 15, including: input unit 16, such as keyboard, mouse, etc.; output unit 17, such as various types of displays, speakers, etc.; storage unit 18, such as disk, optical disk, etc.; and communication unit 19, such as network card, modem, wireless transceiver, etc. Communication unit 19 allows electronic device 10 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0141] Processor 11 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. Processor 11 performs the various methods and processes described above, such as a question-answering method based on retrieval enhancement generation.

[0142] In some embodiments, the search-enhanced question-answering method can be implemented as a computer program tangibly contained in a computer-readable storage medium, such as storage unit 18. In some embodiments, part or all of the computer program can be loaded and / or installed on electronic device 10 via ROM 12 and / or communication unit 19. When the computer program is loaded into RAM 13 and executed by processor 11, one or more steps of the search-enhanced question-answering method described above can be performed. Alternatively, in other embodiments, processor 11 can be configured to execute the search-enhanced question-answering method by any other suitable means (e.g., by means of firmware).

[0143] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), system-on-a-chip (SoCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.

[0144] Computer programs used to implement the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when executed by the processor, the computer programs cause the functions / operations specified in the flowcharts and / or block diagrams to be performed. The computer programs may be executed entirely on a machine, partially on a machine, or as a standalone software package, partially on a machine and partially on a remote machine, or entirely on a remote machine or server.

[0145] In the context of this invention, a computer-readable storage medium can be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, apparatus, or device. A computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination thereof. Alternatively, a computer-readable storage medium may be a machine-readable signal medium. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.

[0146] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).

[0147] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or computing systems that include middleware components (e.g., application servers), or computing systems that include frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), blockchain networks, and the Internet.

[0148] A computing system can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a hosting product within the cloud computing service system to address the shortcomings of traditional physical hosts and VPS services, such as high management difficulty and weak business scalability.

[0149] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this invention can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this invention can be achieved, and this is not limited herein.

[0150] The specific embodiments described above do not constitute a limitation on the scope of protection of this invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this invention should be included within the scope of protection of this invention.

Claims

1. A question-answering method based on retrieval enhancement generation, characterized in that, include: The target problem for obtaining user input; Based on the keywords in the target question, keyword retrieval is performed on each text in a pre-established knowledge base; The knowledge base contains financial knowledge texts and internal investment banking texts; each text corresponds to a text vector. Retrieve text vectors similar to the target vector from a pre-established knowledge base, and determine the text corresponding to the retrieved text vectors; the target vector is a vector generated based on the target question. The target question and the retrieval results obtained from the knowledge base are processed based on a large language model to obtain the answer to the target question.

2. The method according to claim 1, characterized in that, The target question and the retrieval results obtained from the knowledge base are processed based on a large language model to obtain the answer to the target question, including: Based on the keyword retrieval score of the first text and the vector similarity score between the first text vector and the target vector, the second text is determined from the first text; the first text is the text retrieved from the knowledge base. The second text is combined with the target question to obtain the target prompt words; The target prompt words are processed based on a large language model to obtain the answer to the target question.

3. The method according to claim 2, characterized in that, Based on the keyword retrieval score of the first text and the vector similarity score between the first text vector and the target vector, the second text is determined from the first text, including: The keyword retrieval score and the vector similarity score are weighted and summed to obtain the total retrieval score for the first text. The first text is sorted according to the total search score, and the third text is determined from the first text based on the sorting results; The second text is determined from the third text based on the degree of matching between the third text and the target question.

4. The method according to claim 3, characterized in that, Based on the degree of matching between the third text and the target question, a second text is determined from the third text, including: The concatenated sequence is processed using a pre-trained cross-encoder to obtain similarity scores between the target question and each third text; the concatenated sequence is obtained by concatenating the target question and the third text. The third text is sorted according to the similarity score, and the second text is determined from each third text based on the sorting result.

5. The method according to claim 1, characterized in that, The update process of the knowledge base includes: Continuously acquire the first type of documents and convert them into fourth type of text; the first type of documents are documents that meet the requirements of high update frequency; The fourth text is processed using a hash algorithm to obtain a unique text identifier; If there is no text in the knowledge base that is the same as the document identifier of the first type of document, then the fourth text and the fourth text vector are inserted into the knowledge base; If a document in the knowledge base has the same document identifier as the first type of document, but the unique text identifier of that document is different from the unique text identifier of the fourth document, then the corresponding document and its text vector in the knowledge base will be replaced with the fourth document and its text vector.

6. The method according to claim 1, characterized in that, The update process of the knowledge base includes: If a second type of document is received, the version of the second type of document and the version of the corresponding text in the knowledge base are obtained; the second type of document is a document that meets the low update frequency requirement. If the version of the second type of document is a newer version, then the second type of document is converted into text, and the text and the corresponding text vector are saved to the knowledge base.

7. The method according to claim 1, characterized in that, After obtaining the target question input by the user, the method further includes: Based on a pre-established financial security rule base, the target question is subjected to security verification; if the security verification result meets the security requirements, then the step of performing keyword retrieval of each text in the pre-established knowledge base based on the keywords in the target question is executed. After obtaining the answer to the target question, the method further includes: The answer to the target question is processed based on a pre-trained compliance check model. If the output result shows that the answer is compliant, the answer is processed for sensitive information, and the processed answer is fed back to the user.

8. A question-answering device based on retrieval enhancement generation, characterized in that, include: The target question acquisition module is used to acquire the target question input by the user. The keyword retrieval module is used to perform keyword retrieval on each text in a pre-established knowledge base based on keywords in the target question; the knowledge base stores financial knowledge texts and internal investment banking texts; each text corresponds to a text vector; The vector retrieval module is used to retrieve text vectors similar to the target vector from a pre-established knowledge base and determine the text corresponding to the retrieved text vectors; the target vector is a vector generated based on the target question. The answer generation module is used to process the target question and the retrieval results obtained from the knowledge base based on the large language model to obtain the answer corresponding to the target question.

9. An electronic device, characterized in that, The electronic device includes: At least one processor; and A memory communicatively connected to the at least one processor; wherein, The memory stores a computer program that can be executed by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform the question-answering method based on retrieval enhancement generation as described in any one of claims 1-7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions that cause a processor to execute the question-answering method based on retrieval enhancement generation as described in any one of claims 1-7.

Citation Information

Cited By

  • Retrieval enhancement generation method and device, equipment, storage medium and program product

    CN121501967A

  • Search enhancement generation method, apparatus, device, storage medium, and program product

    CN121501967B

  • Retrieval method, device and equipment based on large model, medium and product

    CN121705395A