Text question and answer method and system based on hybrid retrieval enhancement generation
By generating multidimensional dense vectors and using a hybrid retrieval technique with keywords, combined with a large language model, the timeliness and lack of professional knowledge in text question answering systems are solved, achieving efficient and accurate text question answering.
Patent Information
- Application Number
- CN202511184413.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-22
- Publication Date
- 2025-11-07
AI Technical Summary
Text question answering systems based on large language models suffer from time window limitations, cannot dynamically acquire the latest information, lack vertical domain expertise, resulting in low credibility when answering professional questions. Furthermore, single-vector retrieval in RAG technology is susceptible to semantic drift, leading to low retrieval accuracy, weak multimodal support, and poor timeliness.
By generating multidimensional dense vectors and keyword retrieval data, and combining them with hybrid retrieval weights to perform hybrid retrieval, multi-source retrieval results are obtained. These results are then fused using a large language model to generate structured answer data, employing dynamic weight adjustment and multimodal fusion techniques.
It improves the timeliness and retrieval accuracy of the text question-answering process, supports multimodal information processing, enhances the ability to integrate knowledge from vertical domains, and ensures the accuracy and credibility of the answers.
Smart Images

Figure CN120910219A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of artificial intelligence, in particular to a text question and answer method and system based on hybrid retrieval augmented generation. BACKGROUND
[0002] Text question and answer is a human-computer interaction mode based on natural language processing, which can automatically answer questions in the form of text input by users through a question and answer system. In the text question and answer process, the user inputs a question sentence in the form of natural language text to a large language model, and the large language model identifies the semantic information of the question sentence through natural language recognition, and obtains information related to the question from a knowledge base or text resources through retrieval. Based on natural language generation rules, the answer sentence is generated according to the relevant information obtained by retrieval and output to the user.
[0003] The text question and answer system based on the large language model has a time window limit and cannot dynamically obtain the latest information, that is, the knowledge timeliness of the question and answer system is insufficient. In addition, the general large model lacks vertical field professional knowledge, such as enterprise internal documents, industry technical standards, etc., resulting in low credibility when answering professional questions. In addition, the large language model is prone to fabricating facts based on the characteristics of generating text based on probability, and there is a risk of illusion. Enterprise sensitive data cannot be directly input into the public cloud large language model, making it difficult for the question and answer system to realize localized knowledge fusion and permission control.
[0004] In order to improve the problems existing in the above question and answer process, part of the question and answer system can also combine the large language model with the retrieval augmented generation (RAG) technology, retrieve relevant information from the external knowledge base to enhance the output of the large language model, and generate more accurate and rich context responses. However, the single vector retrieval in the RAG technology is easily affected by semantic drift, resulting in low retrieval accuracy, and the ability to process context is insufficient, the multi-modal support ability is weak, and the timeliness is poor. SUMMARY
[0005] Therefore, the embodiments of the present application provide a text question and answer method and system based on hybrid retrieval augmented generation to solve the problem of poor timeliness in the text question and answer process.
[0006] According to one aspect of the present application, a text question and answer method based on hybrid retrieval augmented generation is provided, the method comprising:
[0007] obtaining a question sentence, the question sentence being a natural language text input by a user;
[0008] generate retrieval data based on the question sentence, the retrieval data including a multi-dimensional dense vector generated according to the question sentence and a keyword extracted from the question sentence;
[0009] perform hybrid retrieval based on the hybrid retrieval weight and the retrieval data, to obtain multi-source retrieval result information, the multi-source retrieval result information including a business rule paragraph obtained based on matching of the multi-dimensional dense vector and a target document obtained based on retrieval in a document library based on the keyword;
[0010] fuse the multi-source retrieval result information by a large language model, to generate structured answer data.
[0011] In some embodiments, generating retrieval data based on the question sentence includes:
[0012] obtain a noise filtering rule, the noise rule including a regular expression constructed according to a business document writing manner in a current business field;
[0013] determine noise text in the question sentence based on the noise filtering rule;
[0014] delete the noise text in the question sentence, to obtain de-noised data;
[0015] generate the retrieval data according to the de-noised data.
[0016] In some embodiments, generating the retrieval data according to the de-noised data includes:
[0017] invoke a multi-granularity tokenizer, a split benchmark level of the multi-granularity tokenizer including character-based splitting, word-based splitting and phrase-based splitting;
[0018] perform text splitting processing on the de-noised data using the multi-granularity tokenizer, to obtain a tokenization result, the tokenization result including at least one of a character set, a word set and a phrase set; the character set being a tokenization result obtained based on character-based splitting; the word set being a tokenization result obtained based on word-based splitting; the phrase set being a tokenization result obtained based on phrase-based splitting;
[0019] extract the keyword from the tokenization result, and generate the multi-dimensional dense vector based on the tokenization result.
[0020] In some embodiments, generating the multi-dimensional dense vector based on the tokenization result includes:
[0021] obtain a pre-trained vector embedding model, the vector embedding model including a multi-layer encoder configured to generate a continuous vector representation containing context information according to an input word sequence;
[0022] inputting the word segmentation result into the vector embedding model to extract semantic information of the word segmentation result by the vector embedding model;
[0023] outputting the multi-dimensional dense vector according to the semantic information, the multi-dimensional dense vector being a target position vector extracted from a hidden state of the vector embedding model.
[0024] In some embodiments, the method further comprises:
[0025] obtaining the mixed retrieval weight, the mixed retrieval weight comprising a first weight and a second weight; the first weight being used for performing vector retrieval; the second weight being used for performing keyword retrieval;
[0026] performing vector matching based on the multi-dimensional dense vector to obtain a vector retrieval result, and extracting the business rule paragraph in the vector retrieval result according to the first weight;
[0027] performing keyword retrieval in a document library based on the keyword to obtain a keyword retrieval result, and extracting the target document in the keyword retrieval result according to the second weight;
[0028] combining the business rule paragraph and the target document to generate the multi-source retrieval result information.
[0029] In some embodiments, the method further comprises:
[0030] calculating a term frequency of the retrieval data relative to the retrieval result, the term frequency being a ratio of a number of occurrences of the keyword in the target document to a total number of words in the target document;
[0031] calculating an inverse document frequency of the target document in a corpus, the inverse document frequency being obtained by taking a logarithm of a ratio of a total number of documents in the corpus to the number of target documents;
[0032] generating a decay factor according to the term frequency and the inverse document frequency, the decay factor being a product of the term frequency and the inverse document frequency;
[0033] adjusting the mixed retrieval weight based on the decay factor.
[0034] In some embodiments, the method further comprises:
[0035] obtaining a loss function of a retrieval task;
[0036] calculating a loss value of the retrieval task at a continuous time step based on the loss function;
[0037] According to the loss value, a learning speed of the retrieval task is calculated, the learning speed being equal to a ratio of loss values corresponding to two consecutive time steps;
[0038] According to the learning speed, the mixed retrieval weight is dynamically adjusted.
[0039] In some embodiments, according to the mixed retrieval weight and the retrieval data, a mixed retrieval is performed to obtain multi-source retrieval result information, including:
[0040] From the retrieval data, a constraint factor and a rule document are extracted;
[0041] Based on the constraint factor, candidate knowledge data is screened out from a knowledge base, the candidate knowledge data being data in the knowledge base that satisfies the constraint factor;
[0042] According to the rule document, a mixed retrieval is performed in the candidate knowledge data to obtain multi-source retrieval result information.
[0043] In some embodiments, the multi-source retrieval result information is fused by a large language model to generate structured answer data, including:
[0044] The multi-source retrieval result information is input into a large language model;
[0045] Multi-scale features of the multi-source retrieval result information are extracted by a backbone network of the large language model;
[0046] Based on an attention mechanism, a fusion weight of a data source corresponding to the multi-scale features is calculated;
[0047] According to a preset fusion strategy and the fusion weight, feature fusion is performed on the multi-scale features to generate a fusion result, the preset fusion strategy including at least one of a gated channel attention mechanism, a gated spatial attention mechanism, and a multi-scale residual fusion strategy;
[0048] According to the fusion result, structured answer data is generated.
[0049] According to another aspect of the present application, a text question and answer system based on mixed retrieval enhancement generation is provided, the system including:
[0050] An input module is configured to obtain a question sentence, the question sentence being a natural language text input by a user;
[0051] A preprocessing module is configured to generate retrieval data based on the question sentence, the retrieval data including a multi-dimensional dense vector generated according to the question sentence and a keyword extracted from the question sentence;
[0052] The mixed search module is configured to perform mixed search according to the mixed search weight and the search data, and obtain multi-source search result information, which includes business rule paragraphs obtained based on the multi-dimensional dense vector matching and target documents obtained based on the keyword search in the document library.
[0053] The large model generation module is configured to fuse the multi-source search result information by using a large language model, so as to generate structured answer data.
[0054] According to another aspect of the present application, a computer device is provided, which includes a storage medium, a processor, and a computer program stored in the storage medium and executable on the processor, and the processor implements the above-mentioned text question and answer method based on mixed search enhancement when executing the program.
[0055] According to another aspect of the present application, a storage medium is provided, which stores a computer program, and the program is executed by a processor to implement the above-mentioned text question and answer method based on mixed search enhancement.
[0056] According to the above technical solution, the present application provides a text question and answer method and system based on mixed search enhancement, which first generates search data including a multi-dimensional dense vector and a keyword based on a question statement after obtaining the question statement. Then, mixed search is performed according to the mixed search weight and the search data to obtain multi-source search result information. The multi-source search result information includes business rule paragraphs obtained based on the multi-dimensional dense vector matching and target documents obtained based on the keyword search in the document library. Then, the multi-source search result information is fused by using a large language model to generate structured answer data. The method can perform a multi-path recall mechanism of vector search and keyword search based on the mixed search architecture through the mixed search weight, improve the search efficiency, and improve the timeliness of the text question and answer process. Moreover, the weight of each search channel is optimized in real time through dynamic weight adjustment, and the multi-modal fusion is adopted to support the joint encoding of text and time series data in the text question and answer process, thereby improving the timeliness.
[0057] The above description is only a summary of the technical solutions of the present application. In order to more clearly understand the technical means of the present application, the specific embodiments of the present application can be implemented according to the content of the description, and in order to make the above and other purposes, features and advantages of the present application more obvious and easy to understand, the following specific embodiments of the present application are described. BRIEF DESCRIPTION OF DRAWINGS
[0058] The drawings described herein are used to provide further understanding of the present application, and form a part of the present application. The schematic embodiments of the present application and their descriptions are used to explain the present application, and do not constitute an improper limitation on the present application. In the drawings:
[0059] Figure 1A text question and answer system schematic diagram provided for an embodiment of the present application;
[0060] Figure 2 A text question and answer method flowchart based on hybrid retrieval enhancement generation provided for an embodiment of the present application;
[0061] Figure 3 A text preprocessing flowchart provided for an embodiment of the present application;
[0062] Figure 4 A hybrid retrieval module workflow diagram provided for an embodiment of the present application;
[0063] Figure 5 A system overall architecture diagram provided for an embodiment of the present application;
[0064] Figure 6 A text question and answer system structure schematic diagram based on hybrid retrieval enhancement generation provided for an embodiment of the present application. DETAILED DESCRIPTION
[0065] In the following, the present application will be described in detail with reference to the accompanying drawings and in conjunction with embodiments. It should be noted that the embodiments in the present application and the features in the embodiments can be combined with each other without conflict.
[0066] In an embodiment of the present application, text question and answer is a human-computer interaction method based on natural language processing, which can automatically answer questions in text form proposed by users through a question and answer system. In the text question and answer process, the user inputs a question sentence in the form of natural language text to a large language model, and the large language model identifies the semantic information of the question sentence through natural language recognition, and obtains information related to the question from a knowledge base or a text resource through retrieval. Based on natural language generation rules, the relevant information obtained by retrieval is used to generate an answer sentence, which is output to the user.
[0067] Text question and answer can realize user interaction based on artificial intelligence (AI) robots. Among them, the AI question and answer robot is an intelligent interaction system developed based on artificial intelligence technology. The AI robot has generality, which not only refers to a humanoid appearance structure of a simulation mechanical device, but also refers to an application program with a question and answer function and a combination of an application program and a hardware device. For example, the AI robot is an intelligent assistant installed on a computer, a server, a mobile terminal, a control host and other electronic devices.
[0068] An AI robot can understand various questions raised by a user through natural language processing (NLP) technology, and generate accurate and relevant answer information according to a preset knowledge base, data model, or real-time search, etc. Therefore, in the embodiments of the present application, the process of the AI robot generating and displaying answer information for the question raised by the user is referred to as text question answering.
[0069] As shown in Figure 1 The text question answering process can involve text understanding and text generation based on natural language processing. Text understanding is the core of the AI robot, which can analyze the natural language text input by the user to understand the intent, keywords, and semantic structure thereof. Through analysis and interpretation of the semantics, syntax, context, etc. of the text content, the meaning of the text can be understood. The text understanding can use a text understanding model, which is an artificial intelligence model trained based on sample data with semantic labels. Such as BERT, OpenNLU, etc.
[0070] In some embodiments, when performing text understanding, tokenization can be performed first to divide continuous text strings into meaningful units such as words, phrases, or symbols, etc. Then, part-of-speech tagging is performed on the tokenization results, i.e. determining the part of speech of each word in the sentence, such as noun, verb, adjective, etc. Then, semantic role labeling is performed to identify the semantic roles of each component in the sentence, such as agent, patient, instrument, etc.
[0071] After preprocessing such as tokenization, stem extraction, part-of-speech tagging, named entity recognition, etc., syntax analysis can be performed to analyze the grammatical structure of the sentence, construct a syntax tree, and determine the grammatical relationship between words, such as subject, predicate, object, etc. Based on the word meaning, syntactic structure and semantic role, the overall meaning of the text is inferred. And based on pragmatic analysis algorithm, the actual intent and effect of the text in a specific context is determined.
[0072] After analyzing the natural language text and understanding the user's question, the AI robot can generate a natural and fluent answer. The text generation process needs to use a language generation model, such as a language generation model based on the architecture of deep learning Transformer, etc. The language generation model can learn a large amount of language data, so as to generate answers that conform to the grammatical rules and semantic logic.
[0073] After the text question and answer system is integrated and deployed, a user can input question data through a terminal device after connecting the application to the terminal device through multiple channels such as a webpage, an application, and social media. The text question and answer system can receive the question data input by the user in real time. After receiving the question data input by the user, the text question and answer system inputs the question data into a language understanding model to understand the overall meaning of the question data. Then, the text question and answer system retrieves related information in a knowledge base based on the language understanding result. For example, a retrieval-based question and answer system can find the most relevant answer by matching the user question with the questions in the knowledge base through semantic search technology. After retrieving the related information, the text question and answer system needs to generate answer text. The text question and answer system can generate the answer text based on the retrieved related information, that is, the text question and answer system can directly select the most matching answer data from the knowledge base. The text question and answer system can also obtain the answer text based on the generated text data, that is, the text question and answer system generates the answer data using a deep learning model.
[0074] Since the text question and answer system based on a large language model has a time window limit, it cannot dynamically obtain the latest information, that is, the knowledge of the question and answer system is not timely. In addition, a general large model lacks professional knowledge in a vertical field, such as enterprise internal documents and industry technical standards, which leads to low credibility when answering professional questions. In addition, the large language model is prone to fabricating facts due to the characteristics of generating text based on probability, which has a hallucination risk. Enterprise sensitive data cannot be directly input into a public cloud large language model, which makes it difficult for the question and answer system to realize localized knowledge fusion and permission control.
[0075] To improve the problems in the above question and answer process, in some embodiments, part of the question and answer system can also combine a large language model with retrieval-augmented generation (RAG) technology to enhance the output of the large language model by retrieving related information from an external knowledge base, so as to generate more accurate and more context-rich responses. However, single vector retrieval in the RAG technology is easily affected by semantic drift, which leads to low retrieval accuracy and insufficient ability to process context and weak multi-modal support capability and poor timeliness.
[0076] To solve the problem of poor timeliness in the text question and answer process, the application provides a text question and answer method based on hybrid retrieval-augmented generation in some embodiments. The method can be applied to an electronic device with a data processing function. The electronic device can include but is not limited to a computer, a mobile terminal, a server, a smart wearable device, an industrial control host, and the like. In some embodiments of the application, the text question and answer method based on hybrid retrieval-augmented generation is described by taking an electronic device as an example. It should be understood that the method can also be applied to other types of electronic devices, which will not be described one by one in the embodiments of the application. As shown in the figure, the method includes: Figure 2
[0077] S100: Obtain a question sentence.
[0078] In the process of performing text question answering, the electronic device can first obtain a question sentence, wherein the question sentence is a natural language text input by a user. In some embodiments, the electronic device can provide a user interaction interface, which can include an input control of query information, and the user can input the question sentence through the input control.
[0079] For example, the user can input the question sentence with the content of "the latest ZJ transaction settlement rules" through a text input box. Then the electronic device can obtain the question sentence containing the above natural language text content through the text input box.
[0080] The user can input the question sentence through text input, or can input the question sentence through other interaction methods. That is, in some embodiments, the user can input the question sentence through voice input, handwriting input, image recognition, and the like. For the question sentence input by different interaction methods, the electronic device can set a function module for converting other forms of signals into text information to meet different information input interaction methods, and the embodiments of the present application will not be described one by one.
[0081] In some embodiments, the question sentence can also be obtained based on a business scenario from specific business data. That is, after the user controls the electronic device to display the user interaction interface, the user can upload a business file through a file upload control in the interaction interface. After the business file is uploaded, the electronic device can read the content from the business file, and determine whether the business file contains an operation intention related to text question answering according to the read content. When the business file contains an operation intention related to text question answering, the question sentence can be extracted from the business file according to the operation intention.
[0082] For example, after the user controls the electronic device to display the user interaction interface with the file upload control, the user can upload the ZJ transaction file through the file upload control. Then the electronic device can read data from the ZJ transaction file, and when it is found that the transaction settlement rules position in the transaction file lacks related data, it can be determined that the uploaded ZJ transaction file has an operation intention of transaction settlement rules calculation, and thus a question sentence representing the operation intention can be generated.
[0083] S200: Generate retrieval data based on the question sentence.
[0084] After obtaining the question sentence, the electronic device can generate retrieval data based on the question sentence. In the process of generating the retrieval data, the electronic device can first perform preprocessing such as denoising and word segmentation on the question sentence. That is, for example, Figure 3As shown, in some embodiments, when performing retrieval data generation based on the question sentence, the electronic device can first obtain noise filtering rules, wherein the noise rules include regular expressions constructed according to the writing method of business documents in the current business field. Then, based on the noise filtering rules, determine the noise text in the question sentence, and delete the noise text in the question sentence to obtain de-noised data.
[0085] After obtaining the question sentence, the electronic device can call a text cleaner. The text cleaner can perform noise filtering and other text cleaning on the question sentence based on regular expressions through an internal text cleaning algorithm. During the text cleaning process, the electronic device can perform text processing processes such as defining targets, regular expression writing, applying regular expressions, verifying results, and optimization. That is, before generating retrieval data, the type of noise that needs to be filtered needs to be determined according to the business field to which the question sentence belongs, such as HTML tags, URL links, special characters, numbers, redundant white space characters, and characters in non-target languages.
[0086] Then, according to the type of noise that needs to be filtered in the current business scenario, the corresponding regular expression is written. For example, when filtering noise corresponding to URL link data, a regular expression with content “https?: / / \S+” can be written. The electronic device can match any non-blank character sequence starting with “http: / / ” or “https: / / ” based on the written regular expression, thereby implementing noise filtering on URL link data. For another example, when filtering noise of the type of numbers, a regular expression with content “\d+” can be written to match one or more consecutive numbers through the regular expression.
[0087] After determining the noise data based on the regular expression and removing the noise data from the question sentence, de-noised data can be obtained, and the electronic device can continue to perform data preprocessing based on the de-noised data to generate the retrieval data.
[0088] After performing de-noising processing on the question sentence, the de-noised data can also be subjected to word segmentation processing to obtain the keywords corresponding to the question sentence. In some embodiments, when generating the retrieval data based on the de-noised data, the electronic device can first call a multi-granularity segmenter. The multi-granularity segmenter is a tool that can segment text into different granularities. By supporting multiple segmentation granularities, the multi-granularity segmenter can better handle the characteristics of different languages and text structures, thereby improving the text understanding ability of the model. In order to adapt to natural language texts in different business fields, the segmentation benchmark level of the multi-granularity segmenter includes character-based segmentation, word-based segmentation, and phrase-based segmentation.
[0089] After calling the multi-granularity segmenter, the electronic device can perform text segmentation processing on the de-noised data using the multi-granularity segmenter to obtain a segmentation result. According to different segmentation granularities of the segmenter, the segmentation result obtained by the segmentation processing can include at least one of a character set, a word set, and a phrase set. Obviously, the character set is a segmentation result obtained based on character segmentation; the word set is a segmentation result obtained based on word segmentation; and the phrase set is a segmentation result obtained based on phrase segmentation.
[0090] To adapt to the mixed retrieval function, the retrieval data can include a multi-dimensional dense vector generated according to the question sentence and a keyword extracted from the question sentence. To extract the keyword, after obtaining the segmentation result, the electronic device can extract the keyword from the segmentation result.
[0091] The electronic device can extract the keyword from the segmentation result in a statistical manner, that is, in some embodiments, the electronic device can evaluate the importance of a word in a document based on a statistical method of Term Frequency-Inverse Document Frequency (TF-IDF), and extract the keyword based on the evaluation result. To this end, after obtaining the segmentation result, the electronic device can perform Term Frequency (TF) calculation on individual words in the segmentation result, that is, calculate the term frequency of the retrieval data relative to the retrieval result. Wherein, the term frequency is the ratio of the number of occurrences of the keyword in the target document to the total number of words in the target document. Then calculate the inverse document frequency (Inverse Document Frequency, IDF) of each word in the corpus, and calculate the TF-IDF value of each word based on the inverse document frequency IDF and the term frequency TF, that is, calculate the product of the inverse document frequency IDF and the term frequency TF to obtain the TF-IDF value. Then sort the words according to the TF-IDF value, and select the words with a TF-IDF value higher than a preset score threshold as the keyword.
[0092] In some embodiments, the electronic device can also extract the keyword from the segmentation result based on the word co-occurrence graph method. For example, the electronic device can call a TextRank algorithm based on the segmentation result. TextRank is a ranking algorithm based on a word co-occurrence graph, which extracts keywords and key sentences by constructing a word co-occurrence graph and calculating the weight of each word. To this end, after obtaining the segmentation result, the electronic device can construct a word co-occurrence graph according to the segmentation result, in which the words can be used as nodes and the co-occurrence relationship between the words can be used as edges. Then the weight of each node can be calculated using the PageRank algorithm. Then sort the words according to the weight to select the words with a weight higher than a preset weight threshold as the keyword.
[0093] In addition to extracting keywords from the word segmentation result, the electronic device can also generate the multi-dimensional dense vector based on the word segmentation result. The multi-dimensional dense vector refers to a vector in which each component has a specific numerical value, and most of the components have non-zero values. The numerical value of each dimension of the multi-dimensional dense vector can represent the strength, size, etc. of a certain feature or attribute, so that the multi-dimensional dense vector can carry rich information. In contrast to the dense vector, most of the components of the sparse vector have a value of zero.
[0094] In natural language processing, a neural network model such as Word2Vec or GloVe can be used to map words to a high-dimensional dense vector space to obtain a multi-dimensional dense vector. The neural network model can learn the context information in a large amount of text data, so that words with similar semantics are closer in the vector space.
[0095] In some embodiments, to generate the multi-dimensional dense vector, the electronic device can first obtain a pre-trained vector embedding model when generating the multi-dimensional dense vector based on the word segmentation result. The vector embedding model includes a multi-layer encoder configured to generate a continuous vector representation containing context information according to an input word sequence.
[0096] For example, the pre-trained vector embedding model is a pre-trained language model bge-large-zh based on a BERT-like architecture. The bge-large-zh model can process input text through a multi-layer Transformer encoder. The Transformer encoder processes the input sequence through multiple layers of stacked encoder layers, each of which can include sub-layers such as multi-head self-attention mechanism and feed-forward neural network (FFNN). The multi-head self-attention mechanism can capture long-range dependencies by calculating the attention score of each element in the input sequence with other elements. The feed-forward neural network can further extract features by performing nonlinear transformation on each element. The model is trained on a large amount of Chinese corpus to learn the semantic information of the text.
[0097] After calling the vector embedding model, the electronic device can input the word segmentation result into the vector embedding model to extract the semantic information of the word segmentation result through the vector embedding model, and then output the multi-dimensional dense vector according to the semantic information. The multi-dimensional dense vector is a target position vector extracted from the hidden state of the vector embedding model.
[0098] For example, after the question sentence is processed by the tokenizer into a word segmentation result and a series of word tokens are obtained, the tokenizer will segment the text into subunits suitable for model processing. Then the word tokens are input into the bge-large-zh model for feature encoding. That is, the word segmentation result can be processed by the multi-layer Transformer encoder in the bge-large-zh model. Each layer of the encoder extracts semantic information of the text. After passing through the multi-layer Transformer encoder, the bge-large-zh model can extract embedding vectors, that is, extract the vector at a specific position in the last layer of the hidden state of the bge-large-zh model as the embedding representation of the text. The last layer hidden state of the [CLS] token can be used as the embedding of the entire text.
[0099] S300: Perform hybrid retrieval according to the hybrid retrieval weight and the retrieval data to obtain multi-source retrieval result information.
[0100] After obtaining the retrieval data including the keyword and the dense vector, the electronic device can call the hybrid retrieval module to perform hybrid retrieval, that is, as shown in the figure, Figure 4 The electronic device can perform hybrid retrieval according to the hybrid retrieval weight and the retrieval data to obtain multi-source retrieval result information. The multi-source retrieval result information includes the business rule paragraph obtained based on the multi-dimensional dense vector matching and the target document obtained based on the keyword retrieval in the document library.
[0101] The target document is a business document in the document library that contains the keyword. The business rule paragraph refers to a text paragraph containing the business logic or rules corresponding to the multi-dimensional dense vector. The target document and the business rule paragraph can be derived from the business files in the current business field, such as contract terms, policy explanations, operation guides, user agreements, etc.
[0102] In the hybrid retrieval module, two retrieval channels can be included, namely the vector retrieval channel and the keyword retrieval channel. The vector retrieval channel can use a pre-trained language model such as bge-Large-zh to generate a multi-dimensional dense vector, and perform vector retrieval based on the multi-dimensional dense vector to match and obtain a business rule paragraph.
[0103] The keyword retrieval channel can perform data retrieval in the document library through the keyword to obtain a target document containing the keyword from the document library. In order to obtain more accurate retrieval results, when performing retrieval in the keyword retrieval channel, hybrid retrieval algorithms such as BM42 algorithm and TF-IDF decay factor can also be introduced.
[0104] In some embodiments, to perform the hybrid retrieval, the electronic device can obtain the hybrid retrieval weights before performing the hybrid retrieval according to the hybrid retrieval weights and the retrieval data to obtain the multi-source retrieval result information. The hybrid retrieval weights include a first weight and a second weight. The first weight is used to perform vector retrieval, and the second weight is used to perform keyword retrieval. For example, the initial setting of the hybrid retrieval weights is that the first weight is 60% and the second weight is 40%. Then, the vector retrieval channel can be set to perform vector retrieval according to the 60% weight, and the keyword retrieval channel can be set to perform keyword retrieval according to the 40% weight to form a multi-path recall mechanism based on a higher recall rate.
[0105] Based on the multi-dimensional dense vector, vector matching is performed to obtain a vector retrieval result, and the target document is extracted from the vector retrieval result according to the first weight. Meanwhile, keyword retrieval is performed in the document library based on the keyword to obtain a keyword retrieval result, and the target document is extracted from the keyword retrieval result according to the second weight. By combining the business rule paragraph and the target document, the multi-source retrieval result information can be generated.
[0106] With the execution of the multi-round dialogue and retrieval process, the electronic device can also dynamically adjust the hybrid retrieval weights based on the acquisition of the multi-source retrieval result information. In some embodiments, the electronic device can dynamically adjust the hybrid retrieval weights based on a TF-IDF decay factor. The TF-IDF decay factor is a parameter used to adjust the TF-IDF weight calculation, which can optimize the effect of text feature extraction when processing long texts or specific types of documents.
[0107] To this end, the electronic device can first calculate the term frequency of the retrieval data relative to the retrieval result. The term frequency is the ratio of the number of occurrences of the keyword in the target document to the total number of words in the target document. For example, for a keyword t and a document d, the electronic device can first count the number of occurrences of the keyword t in the document d, and then count the total number of words in the document d. By calculating the ratio of the number of occurrences of the keyword in the target document to the total number of words in the target document, the term frequency is obtained. That is, the term frequency TF(t, d) = the number of occurrences of the keyword t in the document d / the total number of words in the document d.
[0108] The electronic device can also calculate the inverse document frequency of the target document in the corpus while calculating the term frequency. The inverse document frequency is obtained by taking the logarithm of the ratio of the total number of documents in the corpus and the number of target documents. For example, after retrieving the target document containing the keyword t, the number of target documents containing the keyword t Nt is counted first, and then the total number of documents N in the corpus is counted. Then, the ratio of the total number of documents N and the number of documents Nt is calculated, and the logarithm of the ratio is taken to determine the inverse document frequency, i.e., IDF(t) = log(N / Nt).
[0109] The decay factor is generated according to the term frequency and the inverse document frequency. The decay factor is the product of the term frequency and the inverse document frequency, so the mixed retrieval weight is adjusted based on the decay factor. For example, the final value of TF-IDF(t, d) is the product of TF(t, d) and IDF(t), i.e., TF-IDF(t, d) = TF(t, d) x IDF(t). The decay factor is a parameter between 0 and 1 (such as decay) for attenuating the TF value. As the term frequency increases, the growth rate of the TF value gradually slows down, thereby avoiding the excessive influence of some high-frequency words on the TF-IDF value.
[0110] In some embodiments, the electronic device can optimize the weight of each retrieval channel in real time based on a dynamic weight adjustment (DWA) module of the attention mechanism to achieve dynamic adjustment of the mixed retrieval channel. To this end, the electronic device can first obtain the loss function of the retrieval task, and then calculate the loss value of the retrieval task at consecutive time steps based on the loss function. Then, the learning rate of the retrieval task is calculated according to the loss value. The learning rate is equal to the ratio of the loss values at consecutive time steps. Then, the mixed retrieval weight is dynamically adjusted according to the learning rate.
[0111] For example, during the execution of the mixed retrieval process, the electronic device can call the DWA module. The DWA module can define the task and the loss function according to the current business scenario, i.e., in a multi-task learning scenario, it can include the loss of the classification task Lcls and the loss of the positioning task Lloc. After defining the task and the loss function, the electronic device can initialize the weight, i.e., initialize the weight for each task. These weights can be learnable parameters, such as using nn.Parameter of PyTorch to define. Then, the learning rate of the task is calculated, i.e., the learning rate is calculated according to the loss change of the task at consecutive time steps. For task k, the learning rate at time step t-1 can be calculated by the loss values at consecutive time steps, i.e., k w k (t-1) = L k(t-1). Wherein, L k (t-1) and L k (t-2) are the loss values of task k at time steps t-1 and t-2, respectively.
[0112] After calculating the learning rate, the electronic device can dynamically adjust the weight of the hybrid retrieval based on the learning rate. The adjustment of the weight can be realized through the following formula:
[0113]
[0114] Where, w k (t-1) is the learning rate of task k at time step t-1; w i (t-1) is the learning rate of task i at time step t-1; T is a temperature coefficient for adjusting the difference between different task weights; K is a coefficient factor to ensure that the sum of the weights is K, i.e., to normalize the weights.
[0115] In some embodiments, when performing hybrid retrieval, the electronic device can also perform multi-level filtering on the document library through a constraint factor to reduce the amount of data matching in the retrieval process. That is, when performing hybrid retrieval based on the hybrid retrieval weight and the retrieval data to obtain multi-source retrieval result information, the electronic device can first extract a constraint factor and a rule document from the retrieval data, and then filter out candidate knowledge data from the knowledge base based on the constraint factor. Wherein, the candidate knowledge data is the data in the knowledge base that meets the constraint factor. Then perform hybrid retrieval in the candidate knowledge data according to the rule document to obtain multi-source retrieval result information.
[0116] For example, after the electronic device obtains the user input content "the latest ZJ transaction settlement rules" question sentence, after filtering noise and word segmentation processing by the text cleaner and multi-granularity tokenizer, the pre-processing module extracts the time factor "latest" and the rule document "ZJ transaction settlement rules". When performing hybrid retrieval, the candidate knowledge data that meets the "latest" time constraint can be filtered out from the document library based on the time factor "latest". Then perform hybrid retrieval in the candidate knowledge data according to the rule document "ZJ transaction settlement rules" to obtain ZJ transaction settlement data.
[0117] S400: Fuse the multi-source retrieval result information through the large language model to generate structured answer data.
[0118] After obtaining the multi-source retrieval result information through hybrid retrieval, the electronic device can input the multi-source retrieval result information into the large language model, and fuse the multi-source retrieval result information through the large language model to generate structured answer data.
[0119] For example, the electronic device can generate an answer sentence by using a large language model such as LLaMA-3.1 70b, that is, after matching the business rule paragraph A through vector retrieval by the hybrid retrieval module and obtaining the target document B of the document library based on keyword retrieval, the electronic device can input the business rule paragraph A and the target document B into the LLaMA-3.1 70b large language model. The LLaMA-3.1 70b large language model extracts the content related to the transaction settlement details from the business rule paragraph A and the target document B, and generates structured answer data with the content of “ZJ transaction settlement details are AAAA, BBBB” based on a natural language generation algorithm, and displays the answer data in the form of an answer sentence in the user interaction interface.
[0120] The large language model can also dynamically fuse multi-source retrieval results based on an attention gate mechanism to generate structured answer data. That is, in some embodiments, when the electronic device fuses the multi-source retrieval result information through the large language model to generate structured answer data, the electronic device can first input the multi-source retrieval result information into the large language model, and extract multi-scale features of the multi-source retrieval result information through the backbone network of the large language model. Then, based on the attention mechanism, the fusion weight of the multi-scale features corresponding to the data source is calculated, and the multi-scale features are executed according to the preset fusion strategy and the fusion weight to generate a fusion result. The preset fusion strategy includes at least one of a gate channel attention mechanism, a gate spatial attention mechanism, and a multi-scale residual fusion strategy. Then, the structured answer data is generated according to the fusion result.
[0121] For example, when generating structured answer data, the large language model can first perform feature extraction, that is, by multi-source data input, input data from different sources into the large language model. Then use the backbone network such as Swin Transformer to extract multi-scale features. The backbone network will gradually reduce the resolution of the input feature map and expand the receptive field to extract features at different scales. Based on the attention gate channel mechanism (AGCA), the large language model can use the attention gate channel mechanism (AGCA) to fuse the features of the two sources at the encoding part of the feature fusion. AGCA improves the correlation between the features of the two sources through a competition / cooperation mechanism, and extracts complementary features of the two sources. Based on the attention gate space mechanism (AGSA), the large language model can use the attention gate space mechanism (AGSA) to dynamically filter out some high-level semantic features that affect small target detail features at the decoding part. And use the spatial context information to filter out accurate detail features. And based on the multi-scale residual fusion strategy (MFRF), the large language model can fully capture multi-scale context information through the multi-scale feature residual fusion strategy to strengthen the attention of detail features. Calculate the weight of each source through the attention mechanism, and dynamically adjust the weight according to the requirements of the current task. The features of different sources are weighted and fused according to the calculated weight to obtain the final fusion result.
[0122] In some embodiments, in order to obtain more accurate and integrate fact checking modules at the output layer, the fact checking module is a key component to ensure the accuracy and reliability of the model generated content. The fact checking module can provide fact checking based on the prompt chain, multi-modal fact checking, enhanced fact detection framework, and improve the output accuracy of the large language model.
[0123] By applying the technical solutions of the above embodiments, the above embodiments provide a text question and answer method based on hybrid retrieval enhanced generation. After obtaining the question statement, the method first generates retrieval data including multi-dimensional dense vectors and keywords based on the question statement. Then, according to the hybrid retrieval weight and the retrieval data, hybrid retrieval is performed to obtain multi-source retrieval result information. Among them, the multi-source retrieval result information includes the business rule paragraph obtained based on the multi-dimensional dense vector matching and the target document obtained by keyword retrieval in the document library. Then, the large language model fuses the multi-source retrieval result information to generate structured answer data. The method can perform multi-path recall mechanism of vector retrieval and keyword retrieval through hybrid retrieval weight based on hybrid retrieval architecture, improve retrieval efficiency, and improve the timeliness problem of the text question and answer process. And through dynamic weight adjustment to optimize the weight of each retrieval channel in real time, and using multi-modal fusion, the text question and answer process supports the joint encoding of text and time series data, and improves the timeliness.
[0124] In some embodiments, as a specific implementation of the text question answering method based on hybrid retrieval enhancement generation described in the above embodiments, the present application also provides a text question answering system based on hybrid retrieval enhancement generation, as shown in Figure 5 、 Figure 6 The system comprises:
[0125] An input module configured to obtain a question sentence, the question sentence being a natural language text input by a user;
[0126] A preprocessing module configured to generate retrieval data based on the question sentence, the retrieval data comprising a multi-dimensional dense vector generated according to the question sentence and a keyword extracted from the question sentence;
[0127] A hybrid retrieval module configured to perform hybrid retrieval according to a hybrid retrieval weight and the retrieval data to obtain multi-source retrieval result information, the multi-source retrieval result information comprising a business rule paragraph obtained based on matching of the multi-dimensional dense vector and a target document obtained based on retrieval in a document library based on the keyword;
[0128] A large model generation module configured to fuse the multi-source retrieval result information through a large language model to generate structured answer data.
[0129] It should be noted that other corresponding descriptions of the various functional units involved in the text question answering system based on hybrid retrieval enhancement generation provided by the embodiments of the present application can refer to the corresponding descriptions in the text question answering method based on hybrid retrieval enhancement generation provided by the above embodiments, which will not be described here.
[0130] The embodiments of the present application also provide a computer device, which can be a personal computer, a server, a network device, etc. The computer device comprises a bus, a processor, a memory, and a communication interface, and can further comprise an input / output interface and a display device. The processor of the computer device is configured to provide computing and control capabilities. The memory of the computer device comprises a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operating system and the computer program in the non-volatile storage medium to run. The database of the computer device is configured to store location information. The network interface of the computer device is configured to communicate with an external terminal through a network connection. The computer program is executed by the processor to implement the steps in the method embodiments.
[0131] Those skilled in the art can understand that the structure of the computer device described above is only part of the structure related to the scheme of the present application, and does not constitute a limitation on the computer device to which the scheme of the present application is applied. The specific computer device can include more or fewer components, or combine certain components, or have a different arrangement of components.
[0132] In one embodiment, a computer readable storage medium is also provided, which can be non-volatile or volatile, and has stored thereon a computer program. The computer program is executed by a processor to implement the steps in the above method embodiments.
[0133] In one embodiment, a computer program product is also provided, which includes a computer program. The computer program is executed by a processor to implement the steps in the above method embodiments.
[0134] It should be noted that the user information (including but not limited to user equipment information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in the present application are all information and data authorized by the user or authorized by all parties.
[0135] It can be understood by those skilled in the art that all or part of the processes in the above embodiments can be completed by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer readable storage medium. When the computer program is executed, it can include the processes of the above embodiments.
[0136] In the embodiments provided in the present application, any reference to memory, database or other medium can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical storage, high-density embedded non-volatile memory, resistive memory (ReRAM), magnetoresistive memory (MRAM), ferroelectric memory (FRAM), phase change memory (PCM), graphene memory, etc.
[0137] Volatile memory can include random access memory (RAM) or external cache memory, etc. As an illustration but not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc.
[0138] The database involved in each of the embodiments provided in the present application can include at least one of a relational database and a non-relational database. The non-relational database can include a distributed database based on a blockchain, and the like, without being limited thereto. The processor involved in each of the embodiments provided in the present application can be a general-purpose processor, a graphics processor, a digital signal processor, a programmable logic device, a data processing logic device based on quantum computing, and the like, without being limited thereto.
[0139] Each of the technical features of the above embodiments can be combined arbitrarily. In order to make the description simple, each of the technical features in the above embodiments is not described in all possible combinations, however, as long as the combination of the technical features does not exist contradictory, it should be considered as the scope of the present disclosure.
[0140] The above-described embodiments only express several implementation manners of the present application, and the description is relatively specific and detailed, but it should not be understood as a limitation on the scope of the patent of the present application. It should be pointed out that, for ordinary skilled persons in the art, several modifications and improvements can be made without departing from the concept of the present application, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the appended claims.
Claims
1. A method for enhancing generated text question answering based on hybrid retrieval, characterized in that, The method comprises: acquiring a question sentence, the question sentence being natural language text input by a user; generating retrieval data based on the question sentence, the retrieval data comprising a multi-dimensional dense vector generated according to the question sentence and a keyword extracted from the question sentence; performing hybrid retrieval according to a hybrid retrieval weight and the retrieval data to obtain multi-source retrieval result information, the multi-source retrieval result information comprising a business rule paragraph obtained based on matching of the multi-dimensional dense vector and a target document obtained by retrieval in a document library based on the keyword; fusing the multi-source retrieval result information by a large language model to generate structured answer data.
2. The method of claim 1, wherein, Generating retrieval data based on the question sentence comprises: acquiring a noise filtering rule, the noise rule comprising a regular expression constructed according to a business document writing manner in a current business field; determining noise text in the question sentence based on the noise filtering rule; deleting the noise text in the question sentence to obtain de-noised data; generating the retrieval data according to the de-noised data.
3. The method of claim 2, wherein, Generating the retrieval data according to the de-noised data comprises: calling a multi-granularity tokenizer, the tokenization benchmark level of the multi-granularity tokenizer comprising character-based tokenization, word-based tokenization and phrase-based tokenization; performing text tokenization processing on the de-noised data using the multi-granularity tokenizer to obtain tokenization results, the tokenization results comprising at least one of a character set, a word set and a phrase set; the character set being tokenization results obtained based on character tokenization; the word set being tokenization results obtained based on word tokenization; the phrase set being tokenization results obtained based on phrase tokenization; extracting the keyword from the tokenization results and generating the multi-dimensional dense vector based on the tokenization results.
4. The method of claim 3, wherein, Generating the multi-dimensional dense vector based on the tokenization results comprises: acquiring a pre-trained vector embedding model, the vector embedding model comprising a multi-layer encoder configured to generate a continuous vector representation containing context information according to an input word sequence; inputting the tokenization results into the vector embedding model to extract semantic information of the tokenization results by the vector embedding model; outputting the multi-dimensional dense vector according to the semantic information, the multi-dimensional dense vector being a target position vector extracted from a hidden state of the vector embedding model.
5. The method of claim 1, wherein, Performing hybrid retrieval according to a hybrid retrieval weight and the retrieval data to obtain multi-source retrieval result information comprises: acquiring the hybrid retrieval weight, the hybrid retrieval weight comprising a first weight and a second weight; the first weight being used for performing vector retrieval; the second weight being used for performing keyword retrieval; performing vector matching based on the multi-dimensional dense vector to obtain vector retrieval results, and extracting the business rule paragraph in the vector retrieval results according to the first weight; performing keyword retrieval in a document library based on the keyword to obtain keyword retrieval results, and extracting the target document in the keyword retrieval results according to the second weight; combining the business rule paragraph and the target document to generate the multi-source retrieval result information.
6. The method of claim 5, wherein, The method further comprises: calculating a term frequency of the search data relative to the search result, the term frequency being a ratio of a number of occurrences of the keyword in the target document to a total number of words in the target document; calculating an inverse document frequency of the target document in a corpus, the inverse document frequency being obtained by taking a logarithm of a ratio of a total number of documents in the corpus to the number of target documents; generating a decay factor according to the term frequency and the inverse document frequency, the decay factor being a product of the term frequency and the inverse document frequency; adjusting the hybrid search weight based on the decay factor.
7. The method of claim 5, wherein, The method further comprises: obtaining a loss function of a search task; calculating a loss value of the search task at a continuous time step based on the loss function; calculating a learning speed of the search task according to the loss value, the learning speed being equal to a ratio of loss values corresponding to two continuous time steps; dynamically adjusting the hybrid search weight according to the learning speed.
8. The method of claim 1, wherein, Performing hybrid search according to the hybrid search weight and the search data to obtain multi-source search result information, comprising: extracting a constraint factor and a rule document from the search data; filtering out candidate knowledge data from a knowledge base based on the constraint factor, the candidate knowledge data being data in the knowledge base that meets the constraint factor; performing hybrid search in the candidate knowledge data according to the rule document to obtain multi-source search result information.
9. The method of claim 1, wherein, Fusing the multi-source search result information through a large language model to generate structured answer data, comprising: inputting the multi-source search result information into a large language model; extracting multi-scale features of the multi-source search result information through a backbone network of the large language model; calculating a fusion weight of a data source corresponding to the multi-scale features based on an attention mechanism; performing feature fusion on the multi-scale features according to a preset fusion strategy and the fusion weight to generate a fusion result, the preset fusion strategy including at least one of a gated channel attention mechanism, a gated spatial attention mechanism, and a multi-scale residual fusion strategy; generating structured answer data according to the fusion result.
10. A text-based question answering system based on hybrid retrieval augmentation enhanced generation, characterized in that, The system comprises: an input module configured to obtain a question sentence, the question sentence being a natural language text input by a user; a preprocessing module configured to generate search data based on the question sentence, the search data including a multi-dimensional dense vector generated according to the question sentence and a keyword extracted from the question sentence; a hybrid search module configured to perform hybrid search according to a hybrid search weight and the search data to obtain multi-source search result information, the multi-source search result information including a business rule paragraph obtained based on matching of the multi-dimensional dense vector and a target document obtained based on search in a document library according to the keyword; a large model generation module configured to fuse the multi-source search result information through a large language model to generate structured answer data.
Citation Information
Cited By
HPC knowledge question and answer method, system and device based on retrieval enhancement generation and term system construction and storage medium
CN122196141A
An HPC trusted data question and answer method and system based on closed-loop feedback, an HPC trusted data knowledge base construction method and device
CN122472177A