Answer acquisition method, computer program product, device and storage medium
By introducing document retrieval modules and attention mechanisms into large language models, dealing with problems and document vector sequences, the problem of large models illusions in answer acquisition is solved, and the accuracy of answer acquisition is improved.
Patent Information
- Application Number
- CN202411989950.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-31
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2044-12-31
AI Technical Summary
As the number of model parameters increases, large models are more likely to have hallucinations in answer acquisition, resulting in a decrease in the accuracy of answer acquisition.
By introducing a document search module into a large language model, after obtaining the problem, the problem vector sequence and multiple target document vector sequences are processed based on the attention mechanism and gating mechanism, the similarity weight and attention weight are calculated, the weighted vector set is generated, and the answer is finally decoded.
It effectively reduces the generation of hallucinations, improves the accuracy of answer acquisition, and enables the model to better understand the content of the problem and target document.
Smart Images

Figure CN119377372B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer technology, and in particular to an answer acquisition method, a computer program product, a device and a storage medium. Background Art
[0002] As the number of model parameters increases, the scores of large models on various benchmarks are gradually approaching or even exceeding the human level. However, although the larger number of parameters improves the model's expressiveness to a certain extent, enabling it to better capture some complex language patterns, it also makes the model more "confident" - even if the generated content is wrong, the model's expression is more convincing, making it difficult for people to distinguish, making the hallucination problem more hidden, resulting in a decrease in the accuracy of answer acquisition. Summary of the invention
[0003] Based on this, it is necessary to provide an answer acquisition method, computer program product, device and storage medium that can improve the accuracy of answer acquisition in response to the above technical problems.
[0004] In order to solve the above technical problems, in a first aspect, a method for obtaining an answer is provided. The method is applied to a target large language model, the target large language model includes a document retrieval module, and the method includes:
[0005] Obtain the question and retrieve multiple target documents related to the question based on the document retrieval module;
[0006] Vectorize the question and multiple target documents to obtain a question vector sequence and multiple target document vector sequences;
[0007] Process the question vector sequence and multiple target document vector sequences based on the attention mechanism, and calculate the similarity weights between the question and each target document; determine the attention weights between the question and each target document based on the gating mechanism; and determine the weighted vector set based on the similarity weights and the attention weights.
[0008] Decode the set of weighted vectors to generate and output the answer corresponding to the question.
[0009] In one embodiment, retrieving multiple target documents related to the question based on the document retrieval module includes:
[0010] Obtain the question and document set input by the user, where the document set includes multiple documents;
[0011] Calculate the similarity value between the question and each document based on a preset method;
[0012] Sort the similarity values between the question and each document;
[0013] According to the similarity value of each document, a preset number of documents are selected as multiple target documents.
[0014] In one embodiment, vectorizing the question and the multiple target documents to obtain a question vector sequence and a multiple target document vector sequence includes:
[0015] Convert the question into a sequence of question vectors;
[0016] Based on the vector retrieval module, multiple target documents are converted into multiple target document vector sequences.
[0017] In one embodiment, processing the question vector sequence and multiple target document vector sequences based on the attention mechanism, and calculating the similarity weights corresponding to the question and each target document includes:
[0018] Process the question vector sequence based on the self-attention mechanism to obtain the relationship between each word code of the question;
[0019] Based on the cross-attention mechanism, the question vector sequence and multiple target document vector sequences are processed to obtain the similarity weight between the question and each target document;
[0020] Among them, processing the question vector sequence and multiple target document vector sequences to obtain the similarity weight between the question and each target document includes: calculating the similarity weight corresponding to the question and each target document based on the similarity weight calculation formula, the question vector sequence and the multiple target document vector sequences.
[0021] In one embodiment, calculating the similarity weights corresponding to the question and each target document based on the similarity weight calculation formula, the question vector sequence, and the plurality of target document vector sequences includes:
[0022] Map the question vector sequence into the query matrix to obtain the query weight matrix;
[0023] Map multiple target document vector sequences into a key matrix to obtain a key weight matrix;
[0024] Map multiple target document vector sequences into a value matrix to obtain a value weight matrix;
[0025] Calculate similarity weights based on a similarity weight calculation formula, a query weight matrix, a key weight matrix, and a value weight matrix;
[0026] The similarity weight calculation formula is as follows:
[0027] ;
[0028] Among them, Q represents the question vector sequence, R represents multiple target document vector sequences, represents the query weight matrix, represents the key weight matrix, represents the value weight matrix, represents the weight matrix that embeds the question into the attention, represents the weight matrix that maps the target document to the attention, represents the weight matrix that maps document values to the output space, T represents the transpose symbol, represents the scaling factor, and softmax represents the probability distribution function.
[0029] In one embodiment, determining the attention weights corresponding to the question and each target document based on the gating mechanism includes:
[0030] Calculate the attention weights corresponding to the question and each target document based on the gating mechanism formula;
[0031] The gating mechanism formula is as follows:
[0032] ;
[0033] in, represents the weight of the gate, Concat represents the concatenation operation, where the concatenation operation refers to concatenating the question vector sequence Q and the i-th target document vector sequence Ri. Represents the target result between the question and the i-th target document information, which is used to analyze the importance of the target document information to answer generation;
[0034] The attention weights are determined based on the target outcome.
[0035] In one embodiment, determining a set of weighted vectors based on a similarity weight and an attention weight includes:
[0036] The key information of the problem of each target document is determined according to the similarity weight corresponding to each target document;
[0037] Learning the problem and target context information of each target document based on the attention weights, and further obtaining the contribution of each target document to the problem based on the target context information;
[0038] Multiply the attention weights by the target document vector sequence corresponding to the attention weights to obtain a weighted target document vector sequence;
[0039] A weighted vector set is generated based on the key information of each target document, the contribution of each target document to the question, the weighted target document vector sequence and the question vector sequence.
[0040] In one of the embodiments, the method further includes: obtaining a target conversation record, and updating the document set based on the target conversation record, wherein the target conversation record includes an answer generated by a target large language model and a historical record of the target large language model obtaining answers based on questions.
[0041] In one embodiment, before obtaining the question, the method further includes:
[0042] The target large language model is trained. The target large language model is trained including:
[0043] Training an input layer of the target large language model based on the first training method to obtain a first loss;
[0044] Based on the second training method, the encoder layer and the decoder layer of the target large language model are jointly trained to obtain a second loss;
[0045] The total loss is calculated based on the first loss and the second loss, and the target large language model is optimized and trained based on the total loss.
[0046] In one embodiment, training the input layer of the target large language model based on the first training method to obtain the first loss includes:
[0047] Obtaining a target question and a search document, the search document comprising a first search document and a second search document, the first search document storing a plurality of search documents related to the target question, and the second search document storing a plurality of search documents unrelated to the target question;
[0048] Calculate a first loss based on a first loss function formula, a first search document, and a second search document;
[0049] The first loss function formula is as follows:
[0050] ;
[0051] in, represents the first loss, represents the similarity between the question and the first retrieved document, Indicates the similarity between the question and the second search document; represents the first retrieved document, represents the second search document, Indicates the target problem.
[0052] In one embodiment, the encoder layer and the decoder layer of the target large language model are jointly trained based on the second training method to obtain the second loss including:
[0053] Input the target question and the retrieval document into the encoder layer and the decoder layer of the target large language model. The encoder layer and the decoder layer process the target question and the retrieval document to generate and output a reference answer;
[0054] Calculate the second loss based on the reference answer and the second loss function calculation formula;
[0055] Among them, the calculation formula of the second loss function is as follows:
[0056] ;
[0057] in, represents the second loss, A represents the reference answer, Indicates the target problem. Retrieve documents. Represents the answer predicted by the encoder layer and the decoder layer at t time steps The probability of, T represents the total length of the reference answer.
[0058] In one embodiment, calculating the total loss based on the first loss and the second loss includes:
[0059] The total loss is calculated based on the first loss, the second loss and the total loss function calculation formula. The total loss function calculation formula is as follows:
[0060] ;
[0061] in, represents the total loss, represents the first loss, represents the second loss, represents the first loss weight, represents the second loss weight.
[0062] In order to solve the above technical problem, in a second aspect, a computer program product is provided, including a computer program, which implements the steps of the method in the first aspect when executed by a processor.
[0063] In order to solve the above technical problem, a third aspect provides a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the following steps when executing the computer program: the processor implements the steps of the above first aspect method when executing the computer program.
[0064] In order to solve the above technical problems, in a fourth aspect, the present application provides a computer-readable storage medium on which a computer program is stored, and when the computer program is executed by a processor, the steps of the method of the above first aspect are implemented.
[0065] Different from the prior art, in this application, by obtaining a question, a document retrieval module based on a target large language model retrieves multiple target documents related to the question; the question and the multiple target documents are vectorized to obtain a question vector sequence and a multiple target document vector sequence; the question vector sequence and the multiple target document vector sequence are processed based on an attention mechanism, and the similarity weights corresponding to the question and each target document are calculated; the attention weights corresponding to the question and each target document are determined based on a gating mechanism; a weighted vector set is determined based on the similarity weight and the attention weight; the weighted vector set is decoded, and the answer corresponding to the question is generated and output. In this way, the document retrieval module is integrated into the large language model, and an attention mechanism is established between the question input by the user and the multiple target documents retrieved, so that the model can more fully understand the content of the question and the target document, fundamentally reducing the generation of hallucinations. The use of this solution can improve the accuracy of answer acquisition. BRIEF DESCRIPTION OF THE DRAWINGS
[0066] Figure 1 A schematic diagram of a flow chart of a method for obtaining an answer in the prior art;
[0067] Figure 2 A schematic diagram of a flow chart of a method for obtaining an answer in one embodiment;
[0068] Figure 3 is a schematic diagram of the structure of a target large language model in one embodiment;
[0069] Figure 4 is a schematic diagram of the structure of the input layer of the target large language model in one embodiment;
[0070] Figure 5 Schematic diagram of the structure of the target large language model encoder layer in one embodiment;
[0071] Figure 6 A schematic diagram of a process of obtaining a dynamic vector library in one embodiment;
[0072] Figure 7 It is a structural block diagram of an answer acquisition device in one embodiment;
[0073] Figure 8 FIG. 4 is a diagram showing the internal structure of a computer device in one embodiment. DETAILED DESCRIPTION
[0074] In order to make the purpose, technical solution and advantages of the present application more clearly understood, the present application is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0075] In the related technology 1, Speculative Retrieval-Augmented Generation (Speculative RAG) is an optimized RAG technology that generates multiple preliminary drafts in parallel, and uses verification models to evaluate and screen these drafts, and finally selects the optimal solution to improve the quality and efficiency of generated content. Compared with the standard retrieval-augmented generation (Standard RAG) method, Speculative RAG not only improves the processing speed, but also enhances the richness and diversity of the output content while ensuring accuracy. This mechanism effectively improves the performance of the model in complex tasks by introducing parallel processing and efficient screening processes.
[0076] However, Speculative RAG needs to generate and evaluate multiple drafts in parallel, which significantly increases the requirements for computing resources. At the same time, this parallel processing and multi-stage evaluation mechanism make the system design more complex. The generation of multiple drafts places more stringent requirements on the network and takes a long time to respond. Multiple drafts can reduce hallucinations to a certain extent, but cannot fundamentally solve the problem of the model ignoring the target document.
[0077] In the related technology 2, Adaptive Retrieval Augmented Generation (Adaptive RAG) can dynamically decide when to retrieve external knowledge and cleverly balance the use of internal knowledge and external knowledge. Based on the characteristics of the input, Adaptive RAG can intelligently determine whether retrieval is needed and automatically select the most appropriate retrieval strategy. The model determines whether a retrieval operation needs to be performed by analyzing the confidence score of the internal state of the language model. At the same time, it uses "honest probe" technology to avoid hallucinations and ensure that the output content is consistent with the real knowledge. This approach not only reduces unnecessary retrieval requests, but also significantly improves the efficiency of the system and the accuracy of the response.
[0078] However, Adaptive RAG analyzes the relationship between user input and search content to a certain extent, making it difficult for the model to ignore search information and reduce the generation of hallucinations, but this strategy that relies on confidence is not necessarily accurate. In addition, the quality of data retrieved by this strategy will affect the accuracy of the model generation results to a certain extent.
[0079] See also Figure 1 , Figure 1Schematic diagram of the working principle of existing retrieval-augmented generation. Retrieval-augmented generation (aka RAG) provides the large language model with information retrieved from certain data sources as the basis for it to generate answers. RAG usually consists of two stages: retrieving context-related information and using the retrieved knowledge to guide the generation process. Basically, RAG uses the information found by the retrieval algorithm as context to help the large model answer user queries. Both the query and the retrieved context are injected into the prompt sent to the large language model (LLM).
[0080] like Figure 1 As shown, first, obtain the question input and the knowledge text, divide the knowledge text into knowledge text blocks, input the knowledge text blocks and the question into the embedding model, and the embedding model converts the knowledge text blocks and the question input into vector representations to obtain the knowledge text vector and the question input vector. A vector database is constructed based on the knowledge text vector and the question input vector, and then the content most relevant to the question input is screened from the vector database, and the most relevant content screened out is combined with the question input to obtain the query retrieval content combination, that is, the prompt word engineering, and the query retrieval content combination is input into the large language model, and the large language model generates the answer and outputs the answer. In the above implementation, more accurate answers are given to the retrieval content and the user's questions provided by the large language model and retrieval enhancement generation, and the dialogue process is added to the memory bank to update the content of the database.
[0081] However, in this implementation, an embedding model is added outside the large language model to achieve text vectorization. The advantage is that the technology of retrieval enhancement generation is more universal and does not require additional training like fine-tuning. The disadvantage is that the large language model may ignore the content provided by the retrieval enhancement generation and still produce hallucinations, the retrieval content generated by the retrieval enhancement is too long to provide key information and instead affects the output quality, and if there is a network delay in the retrieval knowledge base, it will affect the reasoning performance, etc.
[0082] Big model hallucination is a common problem in the current big model development process. The main reason is that the data coverage during training is not wide enough or the weight ratio is low, which leads to the model's reasoning in these areas without basis and fabrication. Retrieval enhancement generation is a key technology to assist model reasoning. It retrieves relevant information based on the question and inputs it into the model to reduce the model's unfounded answers. However, this additional retrieval not only generates network delays, but also interferes with the model's understanding of the original user's question, affecting the accuracy of the model's output results.
[0083] In view of the above technical problems, in one embodiment, Figure 2 As shown, the present application provides an answer acquisition method, which is applied to a target large language model, the target large language model includes a document retrieval module, and the method specifically includes the following steps:
[0084] Step 101: obtain a question, and retrieve multiple target documents related to the question based on a document retrieval module.
[0085] Among them, retrieving multiple target documents related to the question based on the document retrieval module includes: obtaining the question input by the user and a document set, the document set including multiple documents; calculating the similarity value between the question and each document based on a preset method; sorting the similarity value between the question and each document; sorting according to the similarity value of each document, and selecting a preset number of documents as multiple target documents.
[0086] Specifically, a user interface (UI) or application program interface (API) may be provided so that the user can input questions and then obtain the questions input by the user. This interface may be a web form, a chat window, etc. The target document is obtained from a document set, which includes multiple documents. Here, multiple target documents related to the question may be screened out by calculating the similarity between the question and the documents in the document set.
[0087] Exemplarily, feature extraction can be performed on the question and each document in the document set respectively to obtain a question feature vector and a document feature vector of each document in the document set; the similarity between the question feature vector and the document feature vector of each document in the document set can be calculated according to a similarity calculation formula to obtain a similarity value.
[0088] The similarity calculation formula can be shown as follows:
[0089] ;
[0090] in, Represents the problem feature vector The document feature vector of any document in the document set Similarity value of ; Represents the user feature vector The document feature vector of any document in the document set The dot product of Represents the user feature vector The document feature vector of any document in the document set The Euclidean norm of .
[0091] After obtaining the similarity value calculated between the user feature vector and the document feature vector of each document in the document set, the similarity values are arranged from large to small, and the documents are selected from high to low according to the similarity value of each document. K number of documents can be selected as target documents, where the value of K can be set according to actual needs, and the specific value of K is not limited here.
[0092] In a feasible implementation, a question can be obtained, and a graph containing all documents related to the question and document associations can be constructed based on the question. The graph is used as the basic unit of neural network processing, and an initial feature vector is assigned to each node in the graph; the graph and the initial feature vector corresponding to each node are input into a neural network model of an attention mechanism for multiple rounds of reasoning to obtain a final context-rich representation of each node; several target nodes most relevant to the question are identified from the final representation of all nodes in the graph; and the documents corresponding to each target node are used as multiple target documents corresponding to the question.
[0093] After the user submits a question, a series of natural language processing preprocessing steps can be performed on the question. The preprocessing steps here include but are not limited to: word segmentation, part-of-speech tagging, named document recognition, syntactic analysis, etc. Word segmentation can be to divide the question text into independent words or phrases. Part-of-speech tagging can refer to assigning a part of speech (such as noun, verb, adjective, etc.) to each word. Named document recognition can be to identify entities in the text (such as names of people, places, institutions, etc.), which is very important for subsequent knowledge graph queries. Syntactic analysis can be to analyze the structure of a sentence, such as subject, predicate, object, etc., which helps to understand the meaning of the sentence, etc.
[0094] All possible entity names can be extracted from the question by using named entity recognition (NER) technology. These entities may be names of people, places, institutions, concepts, etc., which usually correspond to nodes in the knowledge graph. Match the identified entity names with the entities in the knowledge graph to find the corresponding nodes. This may involve techniques such as fuzzy matching, synonym matching, and spelling correction, because the user input may not be exactly the same as the standard name in the knowledge graph. Analyze the context and semantic information in the question and identify the relationships related to the identified documents. Use the identified documents as nodes and the extracted relationships as edges to build a knowledge graph, which includes all documents and relationships directly related to the question, as well as the indirect connections that may exist between these documents and relationships.
[0095] Use natural language processing technology to parse the question and extract several key elements in the question from the parsing results. Identify documents with specific meanings from the question sentence, such as names of people, places, institutions, time expressions, etc. These documents are often nodes in the knowledge graph and are crucial for subsequent queries and reasoning. Analyze the dependency relationship between words in the sentence and build a dependency syntax tree. This tree structure can clearly represent the grammatical and semantic relationship between the various components in the sentence.
[0096] The graph is used as the basic unit of neural network processing, and an initial feature vector is assigned to each node in the graph. The key element is used as the query node. The query node and the initial feature vector assigned to each node in the graph are input into the neural network model of the attention mechanism for multiple rounds of reasoning. The neural network model here can be a GNN model. In the GNN model, each node will be updated according to its own feature vector and the feature vector of the node adjacent to it (information transmitted through the edge). In each layer of the GNN, the attention mechanism evaluates the importance of the adjacent nodes to the current node and aggregates the information of the adjacent nodes according to these importance weights. This is achieved by calculating the similarity score (such as dot product, cosine similarity, etc.) between the feature vector of the adjacent node and the feature vector of the current node, and then normalizing these scores (such as using the softmax function) to obtain the attention weight. These attention weights are used to weight the information of the aggregated adjacent nodes to update the representation of the current node.
[0097] GNN models usually contain multiple layers, and each time a node passes through a layer, its representation is updated based on the information of its neighbors. During multiple rounds of reasoning, the representation of the node gradually incorporates contextual information from the graph. After each round of reasoning, the node representation becomes richer in contextual information because it contains not only the initial features of the node itself, but also information from its neighbors, neighbors of neighbors, and so on. After multiple rounds of reasoning, the GNN model outputs the final feature representation of each node. The final feature representation of the node is enriched with contextual information from the entire subgraph and can be used for various downstream tasks such as node classification, link prediction, graph classification, etc.
[0098] Calculate the similarity between each final representation and the question, arrange the similarity values from large to small, and select K number of documents corresponding to the final representation with high similarity values as target documents. The value of K here can be set according to actual needs, and the specific value of K is not limited here.
[0099] The target large language model of this application is built based on retrieval enhancement generation technology and large language models. Large language models can be LLMs (Large Language Models) which are a type of deep learning model specifically designed to understand, analyze, and generate human-like text. They learn the patterns and structures of language usage by analyzing large amounts of text data, so that they can perform tasks such as text classification, sentiment analysis, summarization, and translation. The specific structure of the target large language model built based on retrieval enhancement generation technology and large language models in this application is as follows Figure 3 As shown:
[0100] The target large language model consists of an input layer, an encoder layer (graph retrieval enhanced encoder), and a decoder layer (graph dynamic generation decoder).
[0101] Among them, the function of the input layer is to screen multiple target documents related to the question and to vectorize the user input question (illustrated as user input) and multiple target documents (illustrated as retrieval content) and organize them into a form that can be processed by the target large language model. This application describes the process of converting the question and multiple retrieval documents into a vector sequence.
[0102] Step 102, vectorize the question and the multiple target documents to obtain a question vector sequence and a multiple target document vector sequence.
[0103] Specifically, a question is input into an embedding layer in an input layer of a target large language model, and the embedding layer converts the question into a sequence of question vectors; multiple target documents are input into a vector retrieval module in an input layer of a target large language model, and the vector retrieval module converts the multiple target documents into multiple target document vector sequences.
[0104] like Figure 4 As shown, the input layer of the present application includes an embedding layer and a vector retrieval tool FAISS (vector retrieval module). First, the question input by the user is obtained, and the question is input into the embedding layer to obtain a question vector sequence. The target document is a plurality of documents related to the question obtained from the document set (see FIG. , … ), where K represents the total number of target documents.
[0105] After screening multiple target documents related to the question from the document set, the retrieved target documents are converted into vector embedding representations (Ri) through the vector retrieval tool FAISS (Facebook AI Similarity Search), that is, multiple target document vector sequences corresponding to the multiple target documents are obtained, and then the multiple target document vector sequences and the question vector sequences are input into the vector fusion formula. Finally, the multiple target document vector sequences and the question vector sequences are used as the input of the encoder layer and input into the attention mechanism module of the encoder layer.
[0106] The vector fusion formula is as follows:
[0107] ;
[0108] Among them, Input represents the input vector fusion, [CLS] represents the classification tag, Q represents the question, R1; R2...RK represents K target documents, and [SEP] represents the separator tag.
[0109] The document retrieval module in the input layer of the target large language model set in this application retrieves multiple target documents related to the question, converts the question and the target document into a vectorized representation, and converts it into a form that the target large language model can understand, that is, it can be understood that the question and the target document are taken as the global representation of the entire sequence. Here, the question and the target document are fused and converted into a form that the model can understand. The target large language model set in this application receives the search content and the question as input, and the model will not easily ignore the search content; at the same time, adding a more accurate document retrieval module to extract multiple target documents related to the question is conducive to allowing the model to give more accurate answers.
[0110] Step 103, based on the attention mechanism, the question vector sequence and multiple target document vector sequences are processed to calculate the similarity weights corresponding to the question and each target document; the attention weights corresponding to the question and each target document are determined based on the gating mechanism; and the weighted vector set is determined based on the similarity weights and the attention weights.
[0111] Specifically, the question vector sequence and multiple target document vector sequences are processed based on the attention mechanism, and the similarity weights corresponding to the question and each target document are calculated, including: the question vector sequence is input into the self-attention module in the encoder layer of the target large language model, the self-attention module processes the question vector sequence, and obtains the relationship between each word code of the question; the question vector sequence and multiple target document vector sequences are input into the cross-attention module in the encoder layer of the target large language model, the cross-attention module processes the question vector sequence and multiple target document vector sequences, obtains the similarity weights between the question and each target document, and adjusts the model's attention to different document fragments according to the similarity weights.
[0112] The role of the encoder layer in this application is to establish the relationship between the question and the search content according to the requirements and dynamically select the most relevant information. Figure 5 As shown, the encoder layer in the present application includes an attention mechanism module and a gating mechanism module (gate). The attention mechanism module includes a self-attention mechanism module (Self-attention) and a cross-attention mechanism module (Cross-attention). The question vector sequence is input into the attention mechanism module. The question vector sequence is input into the self-attention mechanism to calculate the attention relationship between any question and other questions, so as to help the target large language model understand the relationship between each word code in the question.
[0113] Multiple target document vector sequences are input into the cross-attention mechanism module. The cross-attention mechanism is a special attention mechanism used to process data from different modalities. It is mainly used for information interaction between different inputs, enabling the model to effectively align and focus on contexts from different sources, thereby helping the model better capture the correlation between two inputs. In the cross-attention mechanism, the model uses an input sequence (such as a question) as a query, and then calculates the attention weight associated with it based on another input sequence (such as a text paragraph). This mechanism allows the model to dynamically focus on different inputs and decide which parts are most important. The attention weight is calculated based on the attention weight calculation formula, the question vector, and multiple target document vectors, and more critical information can be screened out based on the attention weight.
[0114] The similarity weights corresponding to the question and each target document are calculated by the cross attention mechanism of the cross attention module. Specifically, the similarity weights corresponding to the question and each target document are calculated based on the similarity weight calculation formula, the question vector sequence, and the multiple target document vector sequences, including:
[0115] Map the question vector sequence into the query matrix to obtain the query weight matrix;
[0116] Map multiple target document vector sequences into a key matrix to obtain a key weight matrix;
[0117] Map multiple target document vector sequences into a value matrix to obtain a value weight matrix;
[0118] Calculate similarity weights based on a similarity weight calculation formula, a query weight matrix, a key weight matrix, and a value weight matrix;
[0119] The similarity weight calculation formula is as follows:
[0120] ;
[0121] Among them, Q represents the question vector sequence, R represents multiple target document vector sequences, represents the query weight matrix, represents the key weight matrix, represents the value weight matrix, represents the weight matrix that embeds the question into the attention, represents the weight matrix that maps the target document to the attention, represents the weight matrix that maps document values to the output space, T represents the transpose symbol, represents the scaling factor, and softmax represents the probability distribution function.
[0122] , , is a trainable projection matrix. The attention mechanism usually involves mapping the input data into three different representation spaces, namely the query space, the key space, and the value space. This is usually achieved by applying three different linear projections, as shown in the following formula:
[0123] ;
[0124] ;
[0125] ;
[0126] Where X represents input data, which is a vector sequence in this application. represents the query matrix, represents the key matrix, Represents a matrix of values.
[0127] The cross attention mechanism focuses on the relationship between the question Q and multiple target documents R, with the aim of learning the similarity between the question and the target document. Specifically, one sequence (the question vector sequence as the query sequence) focuses on another sequence (multiple target document vector sequences as key and value sequences) through the attention mechanism to capture the relationship between the two sequences.
[0128] A gating mechanism module (gate) is added after the cross-attention module to dynamically control the contribution of each target document Ri in generating the answer.
[0129] Specifically, the attention weights corresponding to the question and each target document are determined based on the gating mechanism, including:
[0130] Calculate the attention weights corresponding to the question and each target document based on the gating mechanism formula;
[0131] The gating mechanism formula is as follows:
[0132] ;
[0133] in, represents the weight of the gate, Concat represents the concatenation operation, where the concatenation operation refers to concatenating the question vector sequence Q and the i-th target document vector sequence Ri. Represents the target result between the question and the i-th target document information. The target result is a constant and is used to analyze the importance of the target document information to answer generation. The attention weight is determined based on the target result, and the weighted vector set is determined based on the attention weight.
[0134] Here, the target result is used as the attention weight, and the attention weight Normalization is performed through the gating mechanism to adjust the degree of dependence of the question on each target document. In addition, the entire encoder is composed of multiple such combinations, and each layer can combine the retrieval attention weights generated by the upper layer to optimize the match between the question Q and multiple target documents R layer by layer. By learning an attention weight, i.e., the gating value, based on the gating mechanism, the contribution of different documents in the final generated answer is determined. This allows the model to adaptively adjust its dependence on different documents when answering different questions or retrieving different contexts), that is, the attention weight learns the degree of sharing of each document when generating an answer, and further obtains the contribution of each target document to the question based on the target context information.
[0135] In addition, the gating mechanism generates weights through the sigmoid function. The generated weight values are smooth and continuous, with good transition properties, and are suitable for processing the continuous change relationship between data. The Softmax function, although its name contains "soft", actually outputs a probability distribution, which is displayed as a value between 0 and 1. It takes all categories into consideration during the calculation process, resulting in an abnormally sharp probability distribution of the output, showing a clear peak. When training deepens, the Softmax output tends to be extreme, close to 0 or 1, showing a switching characteristic, that is, the impact on the unselected categories is almost zero.
[0136] The main structure of the gating mechanism includes the forget gate, input gate and output gate. After obtaining the input information of the input gating mechanism module, the forget gate determines which information needs to be retained and which needs to be discarded at the current time. The input gate concentrates on the important information and determines which new information needs to be remembered at the current time step. The input gate saves important information and ignores the unimportant parts of the input information. The output gate determines the information that needs attention and outputs it. The gating mechanism flexibly controls the flow of information in the Shengjing network through the forget gate, input gate and output gate, ensuring that the target large language model can effectively remember important information and filter out irrelevant information, so that it can perform more stably and efficiently when processing long sequence data, which helps the target large language model better understand complex information processing mechanisms.
[0137] In this application, a question vector sequence and multiple target document vector sequences are input into a cross-attention mechanism module to calculate the attention relationship between any question and any target document, so as to help the target large language model understand the relationship between each character code in the target document and the question, and then filter out the key information of the target document, and then combine the gating mechanism to determine the attention weights corresponding to the question and each target document, and determine the weighted vector set according to the attention weights to optimize the matching between the question and the target document, so as to achieve the accuracy of answer acquisition.
[0138] Determining a weighted vector set based on similarity weights and attention weights includes: determining the key information of the problem of each target document according to the similarity weight corresponding to each target document; learning the problem and the target context information of each target document based on the attention weight, and further obtaining the contribution of each target document to the problem based on the target context information; multiplying the attention weight with the target document vector sequence corresponding to the attention weight to obtain a weighted target document vector sequence; generating a weighted vector set based on the key information of the problem of each target document, the contribution of each target document to the problem, the weighted target document vector sequence and the problem vector sequence.
[0139] The dynamic weight vector output by the encoder can be: , where Z is a weighted vector set, Q is a question vector sequence, represents the weighted product of the attention weight corresponding to the first target document and the target document vector sequence corresponding to the first target document, It represents the weighted product of the attention weight corresponding to the second target document and the target document vector sequence corresponding to the second target document. The weighted vector set includes the weighted product of the attention weight corresponding to the target document and the target document vector sequence corresponding to the target document as many target documents as there are target documents. The contribution of the question corresponding to each target document and the key information of the question of the target document are also involved in constructing the weighted vector set, so that the contribution of the question corresponding to each target document and the key information of the question of the target document can be parsed from the weighted vector set later, and the answer to the question can be provided more accurately.
[0140] In this application, the purpose of introducing the cross-attention mechanism is to calculate the relevance between the question and each document, and adjust the model's attention to the fragments in the document according to the relevance (which can be specific to certain key information in the document); the introduction of the gating mechanism is to dynamically fuse the weights and dynamically adjust the contribution of each retrieved document during the training process, thereby determining how the document plays a role in generating answers (equivalent to learning the ranking of the degree of sharing between documents). With the two dimensions of attention mechanism and gating mechanism, one is the relevance to the content of the document, and the other is the ranking of the contribution of the document to the generation of the question, so as to accurately obtain the answer corresponding to the question.
[0141] Step 104: decode the weighted vector set, generate and output the answer corresponding to the question.
[0142] Specifically, the weighted vector set is input into the encoder layer of the target large language model, and the encoder layer calculates the hidden state sequence based on the weighted vector set; calculates the attention score based on the hidden state sequence; updates the hidden state sequence based on the attention score, and then determines the answer corresponding to the question and outputs it.
[0143] The decoder layer in this application includes a dynamic generation decoder. The function of the dynamic generation decoder is to give the final answer based on the output of the encoder. The structure of the dynamic generation decoder is similar to the architecture of the conventional attention mechanism. The dynamic generation decoder decodes the dynamic weighted vector output by the encoder. Obtain the decoded result, input the decoded result into the decoding layer of the dynamic generation decoder and generate the answer token by token.
[0144] Token is the basic unit of text processing in the model, which can be a character code, word, sentence, etc. The encoder here calculates a hidden state for each time step (that is, each token). These hidden states form a hidden state sequence that captures the current token and its context information.
[0145] After the encoder processes the entire input sequence, the output hidden state sequence is usually called encoder_outputs. Each hidden state sequence corresponds to a token in the input sequence. The shape of this sequence is (seq_len, batch_size, hidden_dim), which is the information transmitted between time steps, indicating the final hidden state of the input sequence received at the current time step position. The hidden state of each time step not only considers the embedding vector of the current word, but also combines the information of all previous time steps. That is, the hidden state of the last time step of the encoder is usually used as the initial state of the decoder, which contains the global context information of the entire input sequence.
[0146] The last layer of the decoder is usually a fully connected layer that maps the hidden state to a probability distribution for each word in the vocabulary. The decoder generates a token at each time step until an end token (e.g., "End of Sequence") is generated or the maximum length is reached. This output is compared with the token of the actual target sequence to calculate the loss.
[0147] The decoder's hidden state is initialized with the encoder's final hidden state and takes the input as the first token. The decoder uses the token generated in the previous time step as the input for the current time step. The decoder outputs the predicted token for the current time step. The predicted token is used as the input for the next time step. If it is generated, the decoding is terminated; otherwise, it continues. Output: The final decoder generates a token sequence as the output sequence.
[0148] In the process of dynamically generating a decoder for decoding, the attention score can be used to remove irrelevant content. Here, if the attention score If the value is less than the constant β, this part of the content will be discarded), ensuring the quality of answer generation. Finally, the user and model conversation records are sent to the memory filtering module for unified vectorization processing.
[0149] The formula for the attention score is as follows:
[0150] ;
[0151] in, represents the attention score, softmax represents the probability distribution function, Q represents the question, Represents the i-th target document.
[0152] By receiving the target large language model with retrieval content and question work as input, the target large language model will not easily ignore the retrieval content; at the same time, a more accurate filtering module is added to extract key information, allowing the model to give more accurate results; finally, by maintaining a real-time dynamic memory vector library, the network delay of the retrieval process is greatly reduced.
[0153] In one embodiment, the present application also includes constructing a data set, wherein the data set mainly includes Wikipedia, web page data (Common Crawl), book texts (BookCorpus), C4 (Colossal CleanCrawled Corpus), and various professional literature data (including PubMed, etc.).
[0154] In one embodiment, the present application also provides for obtaining target conversation records, and updating the document set based on the target conversation records. The target conversation records include answers generated by the target large language model and historical records of the target large language model obtaining answers based on questions.
[0155] like Figure 6 As shown, the present application also includes the construction of a real-time dynamic vector library, that is, the document set can be updated in real time in the present application. The document set is a database maintained locally. Its function is to filter the target documents related to the question from the local database as soon as the user inputs the question, and add them to the input and send them to the target large language model for reasoning. The data sources of this vector library include some data sets used in the pre-training stage, the dialogue retrieval records between users and models, and external real-time data obtained after unified vectorization processing by the embedded filtering module. The document set can provide a sequence of document vectors for training when the target large language model is trained, and when obtaining the user input question, according to the target document corresponding to the question, it can keep the vector data updated while reducing the delay caused by network problems in the model when it is in the target document.
[0156] In one embodiment, before obtaining the problem, the method further includes: training the target large language model, and the training of the target large language model includes: training the input layer of the target large language model based on a first training method to obtain a first loss; jointly training the encoder layer and the decoder layer of the target large language model based on a second training method to obtain a second loss; calculating the total loss based on the first loss and the second loss, and optimizing the training of the target large language model based on the total loss.
[0157] Among them, training the input layer of the target large language model based on the first training method to obtain the first loss includes: obtaining a target question and a retrieval document, the retrieval document includes a first retrieval document and a second retrieval document, the first retrieval document stores multiple retrieval documents related to the target question, and the second retrieval document stores multiple retrieval documents unrelated to the target question; calculating the first loss based on the first loss function formula, the first retrieval document and the second retrieval document; the first loss function formula is as follows:
[0158] ;
[0159] in, represents the first loss, represents the similarity between the question and the first retrieved document, Indicates the similarity between the question and the second search document; represents the first search document, represents the second search document, Indicates the target problem.
[0160] This application performs training in the input layer of the target large language model, such as the Sentence-BERT method, in order to enable the retrieval module to quickly find content related to the question. Among them, the input layer mainly processes the target question and retrieves the document, R+ is the document related to the target question (positive sample), and R- is the document unrelated to the question (negative sample). The purpose of using the first loss to train the input layer of the target large language model is to allow the model to learn the matching relationship between the target question input by the user and the relevant document R+, so as to correctly distinguish between positive and negative samples.
[0161] The encoder layer and the decoder layer of the target large language model are jointly trained based on the second training method to obtain the second loss, including: inputting the target question and the retrieval document into the encoder layer and the decoder layer of the target large language model, the encoder layer and the decoder layer process the target question and the retrieval document, generate and output a reference answer; and calculating the second loss based on the reference answer and the second loss function calculation formula; wherein the second loss function calculation formula is as follows:
[0162] ;
[0163] in, represents the second loss, A represents the reference answer, Indicates the target problem. Retrieve documents. Represents the answer predicted by the encoder layer and the decoder layer at t time steps The probability of, T represents the total length of the reference answer.
[0164] In the present application, the encoder of the encoder layer and the decoder of the decoder layer are jointly trained, the target question and the retrieval document are combined and input to the encoder and the decoder, and then the reference answer is generated. The second loss, that is, the standard cross entropy loss is used to train the encoder layer and the decoder layer to optimize the accuracy of the generated answer.
[0165] The total loss is calculated based on the first loss and the second loss, and the total loss includes:
[0166] ;
[0167] in, represents the total loss, represents the first loss, represents the second loss, represents the first loss weight, represents the second loss weight.
[0168] Since the training of the input layer, encoder layer, and decoder layer is independent, the retrieval results may not match the actual needs of generating answers. Finally, the input layer, encoder layer, and decoder layer need to be coordinated, and the entire process from user input of the target question to the retrieval document until the answer A is generated is trained through the total loss. The retrieval quality and generation quality are continuously optimized to improve the accuracy of the model's retrieval answers.
[0169] In this application, a target large language model is constructed based on retrieval enhancement generation technology and a large language model, and the retrieval enhancement process is added to the large language model, so that the large language model can better understand the relationship between the user input question and the retrieved document, so that the model will not ignore the retrieved document information when reasoning and can better understand the question and the retrieved content, and give a more accurate answer to fundamentally reduce the generation of hallucinations. A multi-stage training method is also proposed, which includes training the input layer of the target large language model, and jointly optimizing the encoder of the encoder layer and the decoder of the decoder layer of the target large language model, which can improve the retrieval accuracy of the target large language model. And a local real-time dynamic vector library is constructed, so that the model can be retrieved from this vector library first, which greatly reduces network delay. In addition, through the interaction between the model and the user, the local vector library and the external data are constantly interacting to keep the dynamic vector library updated in real time, which improves the timeliness of the retrieval content.
[0170] It should be understood that although Figure 2 , Figure 6 The steps in the flowchart are shown in sequence as indicated by the arrows, but these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise specified in this document, there is no strict order restriction for the execution of these steps, and these steps can be executed in other orders. Moreover, Figure 2 , Figure 6 At least part of the steps may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily executed at the same time, but can be executed at different times. The execution order of these sub-steps or stages is not necessarily sequential, but can be executed in turn or alternately with other steps or at least part of the sub-steps or stages of other steps.
[0171] In one embodiment, Figure 7 As shown, an answer acquisition device is provided, comprising: an acquisition module 20, a vectorization module 21, a determination module 22 and an output module 23, wherein:
[0172] An acquisition module 20, used for acquiring a question and retrieving a plurality of target documents related to the question based on a document retrieval module;
[0173] A vectorization module 21, used to vectorize the question and the multiple target documents to obtain a question vector sequence and a multiple target document vector sequence;
[0174] A determination module 22 is used to process the question vector sequence and the plurality of target document vector sequences based on the attention mechanism, calculate the similarity weights corresponding to the question and each target document; determine the attention weights corresponding to the question and each target document based on the gating mechanism; and determine a weighted vector set based on the similarity weights and the attention weights;
[0175] The output module 23 is used to decode the weighted vector set, generate and output the answer corresponding to the question.
[0176] In one embodiment, the above device can implement another implementation of the answer acquisition method, and the specific steps are as follows:
[0177] Retrieving multiple target documents related to the question based on the document retrieval module includes:
[0178] Obtain the question and document set input by the user, where the document set includes multiple documents;
[0179] Calculate the similarity value between the question and each document based on a preset method;
[0180] Sort the similarity values between the question and each document;
[0181] According to the similarity value of each document, a preset number of documents are selected as multiple target documents.
[0182] In one embodiment, the above device can implement another implementation of the answer acquisition method, and the specific steps are as follows:
[0183] Vectorize the question and multiple target documents to obtain a question vector sequence and multiple target document vector sequences including:
[0184] Convert the question into a sequence of question vectors;
[0185] Based on the vector retrieval module, multiple target documents are converted into multiple target document vector sequences.
[0186] In one embodiment, the above device can implement another implementation of the answer acquisition method, and the specific steps are as follows:
[0187] Based on the attention mechanism, the question vector sequence and multiple target document vector sequences are processed, and the similarity weights corresponding to the question and each target document are calculated, including:
[0188] Process the question vector sequence based on the self-attention mechanism to obtain the relationship between each word code of the question;
[0189] Based on the cross-attention mechanism, the question vector sequence and multiple target document vector sequences are processed to obtain the similarity weight between the question and each target document;
[0190] Among them, processing the question vector sequence and multiple target document vector sequences to obtain the similarity weight between the question and each target document includes: calculating the similarity weight corresponding to the question and each target document based on the similarity weight calculation formula, the question vector sequence and the multiple target document vector sequences.
[0191] In one embodiment, the above device can implement another implementation manner of the answer acquisition method, and the specific steps are as follows:
[0192] Calculating the similarity weight corresponding to the question and each target document based on the similarity weight calculation formula, the question vector sequence, and multiple target document vector sequences includes:
[0193] Mapping the question vector sequence into the query matrix to obtain the query weight matrix;
[0194] Mapping the multiple target document vector sequences into the key matrix to obtain the key weight matrix;
[0195] Mapping the multiple target document vector sequences into the value matrix to obtain the value weight matrix;
[0196] Calculating the similarity weight based on the similarity weight calculation formula, the query weight matrix, the key weight matrix, and the value weight matrix;
[0197] Among them, the similarity weight calculation formula is as follows:
[0198] ;
[0199] Among them, Q represents the question vector sequence, R represents the multiple target document vector sequences, represents the query weight matrix, represents the key weight matrix, represents the value weight matrix, represents the weight matrix for embedding the question into the attention, represents the weight matrix for mapping the target document to the attention, represents the weight matrix for mapping the document value to the output space, T represents the transpose symbol, represents the scaling factor, and softmax represents the probability distribution function.
[0200] In one embodiment, the above device can implement another implementation manner of the answer acquisition method, and the specific steps are as follows:
[0201] Determining the attention weight corresponding to the question and each target document based on the gating mechanism includes:
[0202] Calculating the attention weight corresponding to the question and each target document based on the gating mechanism formula;
[0203] The gating mechanism formula is as follows:
[0204] ;
[0205] Among them, represents the weight of the gate, Concat represents the concatenation operation, where the concatenation operation refers to concatenating the question vector sequence Q and the i-th target document vector sequence Ri. Represents the target result between the question and the i-th target document information, which is used to analyze the importance of the target document information to answer generation;
[0206] The attention weights are determined based on the target outcome.
[0207] In one embodiment, the above device can implement another implementation of the answer acquisition method, and the specific steps are as follows:
[0208] The weighted vector set determined based on the similarity weight and the attention weight includes:
[0209] The key information of the problem of each target document is determined according to the similarity weight corresponding to each target document;
[0210] Learning the problem and target context information of each target document based on the attention weights, and further obtaining the contribution of each target document to the problem based on the target context information;
[0211] Multiply the attention weights by the target document vector sequence corresponding to the attention weights to obtain a weighted target document vector sequence;
[0212] A weighted vector set is generated based on the key information of each target document, the contribution of each target document to the question, the weighted target document vector sequence and the question vector sequence.
[0213] In one embodiment, the above device can implement another implementation of the answer acquisition method, and the specific steps are as follows:
[0214] The method also includes: obtaining a target conversation record, and updating a document set based on the target conversation record, wherein the target conversation record includes an answer generated by a target large language model and a historical record of the target large language model obtaining answers based on questions.
[0215] In one embodiment, the above device can implement another implementation of the answer acquisition method, and the specific steps are as follows:
[0216] Also include before getting the question:
[0217] The target large language model is trained. The target large language model is trained including:
[0218] Training an input layer of the target large language model based on the first training method to obtain a first loss;
[0219] Based on the second training method, the encoder layer and the decoder layer of the target large language model are jointly trained to obtain a second loss;
[0220] The total loss is calculated based on the first loss and the second loss, and the target large language model is optimized and trained based on the total loss.
[0221] In one embodiment, the above device can implement another implementation of the answer acquisition method, and the specific steps are as follows:
[0222] The input layer of the target large language model is trained based on the first training method, and the first loss includes:
[0223] Obtaining a target question and a search document, the search document comprising a first search document and a second search document, the first search document storing a plurality of search documents related to the target question, and the second search document storing a plurality of search documents unrelated to the target question;
[0224] Calculate a first loss based on a first loss function formula, a first search document, and a second search document;
[0225] The first loss function formula is as follows:
[0226] ;
[0227] in, represents the first loss, represents the similarity between the question and the first retrieved document, Indicates the similarity between the question and the second search document; represents the first search document, represents the second search document, Indicates the target problem.
[0228] In one embodiment, the above device can implement another implementation of the answer acquisition method, and the specific steps are as follows:
[0229] Based on the second training method, the encoder layer and the decoder layer of the target large language model are jointly trained to obtain the second loss including:
[0230] Input the target question and the retrieval document into the encoder layer and the decoder layer of the target large language model. The encoder layer and the decoder layer process the target question and the retrieval document to generate and output a reference answer;
[0231] Calculate the second loss based on the reference answer and the second loss function calculation formula;
[0232] Among them, the calculation formula of the second loss function is as follows:
[0233] ;
[0234] in, represents the second loss, A represents the reference answer, Indicates the target problem. Retrieve documents. Represents the answer predicted by the encoder layer and the decoder layer at t time steps The probability of, T represents the total length of the reference answer.
[0235] In one embodiment, the above device can implement another implementation of the answer acquisition method, and the specific steps are as follows:
[0236] The total loss calculated based on the first loss and the second loss includes:
[0237] The total loss is calculated based on the first loss, the second loss and the total loss function calculation formula. The total loss function calculation formula is as follows:
[0238] ;
[0239] in, represents the total loss, represents the first loss, represents the second loss, represents the first loss weight, represents the second loss weight.
[0240] For the specific definition of the answer acquisition device, please refer to the definition of the answer acquisition method above, which will not be repeated here. Each module in the above-mentioned answer acquisition device can be implemented in whole or in part by software, hardware and a combination thereof. The above-mentioned modules can be embedded in or independent of the processor in the computer device in the form of hardware, or can be stored in the memory of the computer device in the form of software, so that the processor can call and execute the operations corresponding to the above modules.
[0241] In one embodiment, the present application also provides a computer program product, which includes a computer program stored on a non-transitory computer-readable storage medium, and the computer program includes program instructions. When the program instructions are executed by a computer, the computer can execute the answer acquisition methods provided by the above methods.
[0242] In one embodiment, a computer device is provided. The computer device may be a terminal, and its internal structure diagram may be as follows: Figure 8As shown. The computer device includes a processor, a memory, a network interface, a display screen and an input device connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, a method for obtaining an answer is implemented. The display screen of the computer device can be a liquid crystal display screen or an electronic ink display screen, and the input device of the computer device can be a touch layer covering the display screen, or a button, trackball or touchpad set on the computer device housing, or an external keyboard, touchpad or mouse, etc.
[0243] Those skilled in the art will understand that Figure 8 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.
[0244] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the following steps when executing the computer program:
[0245] Obtain the question and retrieve multiple target documents related to the question based on the document retrieval module;
[0246] Vectorize the question and multiple target documents to obtain a question vector sequence and multiple target document vector sequences;
[0247] Process the question vector sequence and multiple target document vector sequences based on the attention mechanism, and calculate the similarity weights between the question and each target document; determine the attention weights between the question and each target document based on the gating mechanism; and determine the weighted vector set based on the similarity weights and the attention weights.
[0248] Decode the set of weighted vectors to generate and output the answer corresponding to the question.
[0249] In one embodiment, when the processor executes the computer program, the processor further implements the following steps:
[0250] Retrieving multiple target documents related to the question based on the document retrieval module includes:
[0251] Obtain the question and document set input by the user, where the document set includes multiple documents;
[0252] Calculate the similarity value between the question and each document based on a preset method;
[0253] Sort the similarity values between the question and each document;
[0254] According to the similarity value of each document, a preset number of documents are selected as multiple target documents.
[0255] In one embodiment, when the processor executes the computer program, the processor further implements the following steps:
[0256] Vectorize the question and multiple target documents to obtain a question vector sequence and multiple target document vector sequences including:
[0257] Convert the question into a sequence of question vectors;
[0258] Based on the vector retrieval module, multiple target documents are converted into multiple target document vector sequences.
[0259] In one embodiment, processing the question vector sequence and multiple target document vector sequences based on the attention mechanism, and calculating the similarity weights corresponding to the question and each target document includes:
[0260] Process the question vector sequence based on the self-attention mechanism to obtain the relationship between each word code of the question;
[0261] Based on the cross-attention mechanism, the question vector sequence and multiple target document vector sequences are processed to obtain the similarity weight between the question and each target document;
[0262] Among them, processing the question vector sequence and multiple target document vector sequences to obtain the similarity weight between the question and each target document includes: calculating the similarity weight corresponding to the question and each target document based on the similarity weight calculation formula, the question vector sequence and the multiple target document vector sequences.
[0263] In one embodiment, when the processor executes the computer program, the processor further implements the following steps:
[0264] The similarity weights corresponding to the question and each target document are calculated based on the similarity weight calculation formula, the question vector sequence, and multiple target document vector sequences, including:
[0265] Map the question vector sequence into the query matrix to obtain the query weight matrix;
[0266] Map multiple target document vector sequences into a key matrix to obtain a key weight matrix;
[0267] Map multiple target document vector sequences into a value matrix to obtain a value weight matrix;
[0268] Calculate similarity weights based on a similarity weight calculation formula, a query weight matrix, a key weight matrix, and a value weight matrix;
[0269] The similarity weight calculation formula is as follows:
[0270] ;
[0271] Among them, Q represents the question vector sequence, R represents multiple target document vector sequences, represents the query weight matrix, represents the key weight matrix, represents the value weight matrix, represents the weight matrix that embeds the question into the attention, represents the weight matrix that maps the target document to the attention, represents the weight matrix that maps document values to the output space, T represents the transpose symbol, represents the scaling factor, and softmax represents the probability distribution function.
[0272] In one embodiment, when the processor executes the computer program, the processor further implements the following steps:
[0273] The attention weights corresponding to the question and each target document are determined based on the gating mechanism:
[0274] Calculate the attention weights corresponding to the question and each target document based on the gating mechanism formula;
[0275] The gating mechanism formula is as follows:
[0276] ;
[0277] in, represents the weight of the gate, Concat represents the concatenation operation, where the concatenation operation refers to concatenating the question vector sequence Q and the i-th target document vector sequence Ri. Represents the target result between the question and the i-th target document information, which is used to analyze the importance of the target document information to answer generation;
[0278] The attention weights are determined based on the target outcome.
[0279] In one embodiment, when the processor executes the computer program, the processor further implements the following steps:
[0280] The weighted vector set determined based on the similarity weight and the attention weight includes:
[0281] The key information of the problem of each target document is determined according to the similarity weight corresponding to each target document;
[0282] Learning the problem and target context information of each target document based on the attention weights, and further obtaining the contribution of each target document to the problem based on the target context information;
[0283] Multiply the attention weights by the target document vector sequence corresponding to the attention weights to obtain a weighted target document vector sequence;
[0284] A weighted vector set is generated based on the key information of each target document, the contribution of each target document to the question, the weighted target document vector sequence and the question vector sequence.
[0285] In one embodiment, when the processor executes the computer program, the processor further implements the following steps:
[0286] The method also includes: obtaining a target conversation record, and updating a document set based on the target conversation record, wherein the target conversation record includes an answer generated by a target large language model and a historical record of the target large language model obtaining answers based on questions.
[0287] In one embodiment, when the processor executes the computer program, the processor further implements the following steps:
[0288] Also include before getting the question:
[0289] The target large language model is trained. The target large language model is trained including:
[0290] Training an input layer of the target large language model based on the first training method to obtain a first loss;
[0291] Based on the second training method, the encoder layer and the decoder layer of the target large language model are jointly trained to obtain a second loss;
[0292] The total loss is calculated based on the first loss and the second loss, and the target large language model is optimized and trained based on the total loss.
[0293] In one embodiment, training the input layer of the target large language model based on the first training method to obtain the first loss includes:
[0294] Obtaining a target question and a search document, the search document comprising a first search document and a second search document, the first search document storing a plurality of search documents related to the target question, and the second search document storing a plurality of search documents unrelated to the target question;
[0295] Calculate a first loss based on a first loss function formula, a first search document, and a second search document;
[0296] The first loss function formula is as follows:
[0297] ;
[0298] in, represents the first loss, represents the similarity between the question and the first retrieved document, Indicates the similarity between the question and the second search document; represents the first search document, represents the second search document, Indicates the target problem.
[0299] In one embodiment, when the processor executes the computer program, the processor further implements the following steps:
[0300] Based on the second training method, the encoder layer and the decoder layer of the target large language model are jointly trained to obtain the second loss including:
[0301] Input the target question and the retrieval document into the encoder layer and the decoder layer of the target large language model. The encoder layer and the decoder layer process the target question and the retrieval document to generate and output a reference answer;
[0302] Calculate the second loss based on the reference answer and the second loss function calculation formula;
[0303] Among them, the calculation formula of the second loss function is as follows:
[0304] ;
[0305] in, represents the second loss, A represents the reference answer, Indicates the target problem. Retrieve documents. Represents the answer predicted by the encoder layer and the decoder layer at t time steps The probability of, T represents the total length of the reference answer.
[0306] In one embodiment, when the processor executes the computer program, the processor further implements the following steps:
[0307] The total loss calculated based on the first loss and the second loss includes:
[0308] The total loss is calculated based on the first loss, the second loss and the total loss function calculation formula. The total loss function calculation formula is as follows:
[0309] ;
[0310] in, represents the total loss, represents the first loss, represents the second loss, represents the first loss weight, represents the second loss weight.
[0311] In one embodiment, a computer readable storage medium is provided, on which a computer program is stored, and when the computer program is executed by a processor, the following steps are implemented:
[0312] Obtain the question and retrieve multiple target documents related to the question based on the document retrieval module;
[0313] Vectorize the question and multiple target documents to obtain a question vector sequence and multiple target document vector sequences;
[0314] Process the question vector sequence and multiple target document vector sequences based on the attention mechanism, and calculate the similarity weights between the question and each target document; determine the attention weights between the question and each target document based on the gating mechanism; and determine the weighted vector set based on the similarity weights and the attention weights.
[0315] Decode the set of weighted vectors to generate and output the answer corresponding to the question.
[0316] In one embodiment, when the computer program is executed by a processor, the following steps are also implemented:
[0317] Retrieving multiple target documents related to the question based on the document retrieval module includes:
[0318] Obtain the question input by the user and the document set, which includes multiple documents;
[0319] Calculate the similarity value between the question and each document based on a preset method;
[0320] Sort the similarity values between the question and each document;
[0321] According to the similarity value of each document, a preset number of documents are selected as multiple target documents.
[0322] In one embodiment, when the computer program is executed by a processor, the following steps are also implemented:
[0323] Vectorize the question and multiple target documents to obtain a question vector sequence and multiple target document vector sequences including:
[0324] Convert the question into a sequence of question vectors;
[0325] Based on the vector retrieval module, multiple target documents are converted into multiple target document vector sequences.
[0326] In one embodiment, when the computer program is executed by a processor, the following steps are also implemented:
[0327] Based on the attention mechanism, the question vector sequence and multiple target document vector sequences are processed, and the similarity weights corresponding to the question and each target document are calculated, including:
[0328] Process the question vector sequence based on the self-attention mechanism to obtain the relationship between each word code of the question;
[0329] Based on the cross-attention mechanism, the question vector sequence and multiple target document vector sequences are processed to obtain the similarity weight between the question and each target document;
[0330] Among them, processing the question vector sequence and multiple target document vector sequences to obtain the similarity weight between the question and each target document includes: calculating the similarity weight corresponding to the question and each target document based on the similarity weight calculation formula, the question vector sequence and the multiple target document vector sequences.
[0331] In one embodiment, when the computer program is executed by a processor, the following steps are also implemented:
[0332] The similarity weights corresponding to the question and each target document are calculated based on the similarity weight calculation formula, the question vector sequence, and multiple target document vector sequences, including:
[0333] Map the question vector sequence into the query matrix to obtain the query weight matrix;
[0334] Map multiple target document vector sequences into a key matrix to obtain a key weight matrix;
[0335] Map multiple target document vector sequences into a value matrix to obtain a value weight matrix;
[0336] Calculate similarity weights based on a similarity weight calculation formula, a query weight matrix, a key weight matrix, and a value weight matrix;
[0337] The similarity weight calculation formula is as follows:
[0338] ;
[0339] Among them, Q represents the question vector sequence, R represents multiple target document vector sequences, represents the query weight matrix, represents the key weight matrix, represents the value weight matrix, represents the weight matrix that embeds the question into the attention, represents the weight matrix that maps the target document to the attention, represents the weight matrix that maps document values to the output space, T represents the transpose symbol, represents the scaling factor, and softmax represents the probability distribution function.
[0340] In one embodiment, when the computer program is executed by a processor, the following steps are also implemented:
[0341] The attention weights corresponding to the question and each target document are determined based on the gating mechanism:
[0342] Calculate the attention weights corresponding to the question and each target document based on the gating mechanism formula;
[0343] The gating mechanism formula is as follows:
[0344] ;
[0345] in, represents the weight of the gate, Concat represents the concatenation operation, where the concatenation operation refers to concatenating the question vector sequence Q and the i-th target document vector sequence Ri. Represents the target result between the question and the i-th target document information, which is used to analyze the importance of the target document information to answer generation;
[0346] The attention weights are determined based on the target outcome.
[0347] In one embodiment, when the computer program is executed by a processor, the following steps are also implemented:
[0348] The weighted vector set determined based on the similarity weight and the attention weight includes:
[0349] The key information of the problem of each target document is determined according to the similarity weight corresponding to each target document;
[0350] Learning the problem and target context information of each target document based on the attention weights, and further obtaining the contribution of each target document to the problem based on the target context information;
[0351] Multiply the attention weights by the target document vector sequence corresponding to the attention weights to obtain a weighted target document vector sequence;
[0352] A weighted vector set is generated based on the key information of each target document, the contribution of each target document to the question, the weighted target document vector sequence and the question vector sequence.
[0353] In one embodiment, when the computer program is executed by a processor, the following steps are also implemented:
[0354] The method also includes: obtaining a target conversation record, and updating a document set based on the target conversation record, wherein the target conversation record includes an answer generated by a target large language model and a historical record of the target large language model obtaining answers based on questions.
[0355] In one embodiment, before obtaining the question, the method further includes:
[0356] The target large language model is trained. The target large language model is trained including:
[0357] Training an input layer of the target large language model based on the first training method to obtain a first loss;
[0358] Based on the second training method, the encoder layer and the decoder layer of the target large language model are jointly trained to obtain a second loss;
[0359] The total loss is calculated based on the first loss and the second loss, and the target large language model is optimized and trained based on the total loss.
[0360] In one embodiment, when the computer program is executed by a processor, the following steps are also implemented:
[0361] The input layer of the target large language model is trained based on the first training method, and the first loss includes:
[0362] Obtaining a target question and a search document, the search document comprising a first search document and a second search document, the first search document storing a plurality of search documents related to the target question, and the second search document storing a plurality of search documents unrelated to the target question;
[0363] Calculate a first loss based on a first loss function formula, a first search document, and a second search document;
[0364] The first loss function formula is as follows:
[0365] ;
[0366] in, represents the first loss, represents the similarity between the question and the first retrieved document, Indicates the similarity between the question and the second search document; represents the first retrieved document, represents the second search document, Indicates the target problem.
[0367] In one embodiment, when the computer program is executed by a processor, the following steps are also implemented:
[0368] Based on the second training method, the encoder layer and the decoder layer of the target large language model are jointly trained to obtain the second loss including:
[0369] Input the target question and the retrieval document into the encoder layer and the decoder layer of the target large language model. The encoder layer and the decoder layer process the target question and the retrieval document to generate and output a reference answer;
[0370] Calculate the second loss based on the reference answer and the second loss function calculation formula;
[0371] Among them, the calculation formula of the second loss function is as follows:
[0372] ;
[0373] in, represents the second loss, A represents the reference answer, Indicates the target problem. Retrieve documents. Represents the answer predicted by the encoder layer and the decoder layer at t time steps The probability of, T represents the total length of the reference answer.
[0374] In one embodiment, when the computer program is executed by a processor, the following steps are also implemented:
[0375] The total loss calculated based on the first loss and the second loss includes:
[0376] The total loss is calculated based on the first loss, the second loss and the total loss function calculation formula. The total loss function calculation formula is as follows:
[0377] ;
[0378] in, represents the total loss, represents the first loss, represents the second loss, represents the first loss weight, represents the second loss weight.
[0379] Those of ordinary skill in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application may include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.
[0380] The technical features of the above embodiments may be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0381] The above embodiments only express several implementation methods of the present application, and the descriptions thereof are relatively specific and detailed, but they cannot be understood as limiting the scope of the invention patent. It should be pointed out that, for ordinary technicians in this field, several variations and improvements can be made without departing from the concept of the present application, and these all belong to the protection scope of the present application.
Claims
1. A method for obtaining an answer, characterized in that: The method is applied to a target large language model, the target large language model includes a document retrieval module, and the method includes: Obtain the question and retrieve multiple target documents related to the question based on the document retrieval module; Vectorizing the question and the multiple target documents to obtain a question vector sequence and a multiple target document vector sequence; processing the question vector sequence and the multiple target document vector sequences based on an attention mechanism, and calculating a similarity weight corresponding to the question and each target document; The step of processing the question vector sequence and the multiple target document vector sequences based on the attention mechanism to calculate the similarity weights corresponding to the question and each target document includes: processing the question vector sequence based on the self-attention mechanism to obtain the relationship between each word code of the question; processing the question vector sequence and the multiple target document vector sequences based on the cross-attention mechanism to obtain the similarity weights between the question and each target document; Determine the attention weights corresponding to the question and each target document based on the gating mechanism; determine a weighted vector set based on the similarity weight and the attention weight; The weighted vector set is decoded based on the attention score to generate and output an answer corresponding to the question, wherein the attention score is used to eliminate content that is not related to the answer.
2. The method according to claim 1, characterized in that The retrieving multiple target documents related to the question based on the document retrieval module includes: Obtaining a question input by a user and a document set, wherein the document set includes multiple documents; Calculate the similarity value between the question and each document based on a preset method; Sort the similarity values between the question and each document; According to the order of the similarity value of each document, a preset number of documents are selected as multiple target documents.
3. The method according to claim 1, characterized in that The step of vectorizing the question and the plurality of target documents to obtain a question vector sequence and a plurality of target document vector sequences comprises: Convert the question into a sequence of question vectors; The multiple target documents are converted into multiple target document vector sequences based on a vector retrieval module.
4. The method according to claim 1, characterized in that The processing of the question vector sequence and the plurality of target document vector sequences to obtain the similarity weight between the question and each target document includes: The similarity weight corresponding to the question and each target document is calculated based on the similarity weight calculation formula, the question vector sequence and the multiple target document vector sequences.
5. The method according to claim 4, characterized in that The calculating of the similarity weights corresponding to the question and each target document based on the similarity weight calculation formula, the question vector sequence and the plurality of target document vector sequences comprises: Map the question vector sequence into the query matrix to obtain the query weight matrix; Map multiple target document vector sequences into a key matrix to obtain a key weight matrix; Map multiple target document vector sequences into a value matrix to obtain a value weight matrix; Calculating similarity weights based on the similarity weight calculation formula, the query weight matrix, the key weight matrix, and the value weight matrix; The similarity weight calculation formula is as follows: Among them, Q represents the question vector sequence, R represents multiple target document vector sequences, and W Q represents the query weight matrix, W K represents the key weight matrix, W V Represents the value weight matrix, QW Q Represents the weight matrix that embeds the question into the attention, RW K represents the weight matrix that maps the target document to the attention, RW V represents the weight matrix that maps document values to the output space, T represents the transpose symbol, represents the scaling factor, and softmax represents the probability distribution function.
6. The method according to claim 1, characterized in that The method of determining the attention weights corresponding to the question and each target document based on the gating mechanism includes: Calculate the attention weights corresponding to the question and each target document based on the gating mechanism formula; The gating mechanism formula is as follows: α i =σ(W gate ·Concat(Q,R i )); Among them, W gate represents the weight of the gate, Concat represents the concatenation operation, where the concatenation operation refers to concatenating the question vector sequence Q and the i-th target document vector sequence Ri, α i Represents the target result between the question and the i-th target document information, which is used to analyze the importance of the target document information to answer generation; The attention weights are determined based on the target outcome.
7. The method according to claim 1, characterized in that The determining of a weighted vector set based on the similarity weight and the attention weight comprises: The key information of the problem of each target document is determined according to the similarity weight corresponding to each target document; Based on the attention weight learning problem and target context information of each target document, further obtaining a contribution of each target document to the problem based on the target context information; Multiplying the attention weight with the target document vector sequence corresponding to the attention weight to obtain a weighted target document vector sequence; A weighted vector set is generated based on the key information of each target document, the contribution of each target document to the question, the weighted target document vector sequence and the question vector sequence.
8. The method according to claim 2, characterized in that: The method also includes: obtaining a target conversation record, and updating a document set based on the target conversation record, wherein the target conversation record includes an answer generated by a target large language model and a historical record of the target large language model obtaining an answer based on a question.
9. The method according to claim 1, characterized in that: The acquisition problem also includes: The target large language model is trained, wherein the training of the target large language model includes: Training an input layer of the target large language model based on the first training method to obtain a first loss; Based on the second training method, the encoder layer and the decoder layer of the target large language model are jointly trained to obtain a second loss; A total loss is calculated based on the first loss and the second loss, and optimization training is performed on the target large language model based on the total loss.
10. The method according to claim 9, characterized in that The step of training the input layer of the target large language model based on the first training method to obtain the first loss includes: Acquire a target question and a search document, wherein the search document includes a first search document and a second search document, wherein the first search document stores a plurality of search documents related to the target question, and the second search document stores a plurality of search documents unrelated to the target question; Calculate a first loss based on a first loss function formula, the first search document, and the second search document; The first loss function formula is as follows: Among them, L retrieval represents the first loss, exp(Q1·R + ) represents the similarity between the question and the first search document, exp(Q1·R - ) represents the similarity between the question and the second search document; R + represents the first search document, R- represents the second search document, and Q1 represents the target question.
11. The method according to claim 9, characterized in that The encoder layer and the decoder layer of the target large language model are jointly trained based on the second training method to obtain the second loss, which includes: Inputting the target question and the retrieval document into the encoder layer and the decoder layer of the target large language model, the encoder layer and the decoder layer process the target question and the retrieval document, generate and output a reference answer; Calculate the second loss based on the reference answer and the second loss function calculation formula; Among them, the calculation formula of the second loss function is as follows: Among them, L generation represents the second loss, A represents the reference answer, Q1 represents the target question, R1 represents the retrieved document, and P gen (A t |Q1, R1) indicates the answer A predicted by the encoder layer and the decoder layer at t time steps t The probability of, T represents the total length of the reference answer.
12. The method according to claim 9, characterized in that The calculating the total loss based on the first loss and the second loss comprises: The total loss is calculated based on the first loss, the second loss and the total loss function calculation formula, and the total loss function calculation formula is as follows: L joint =λ1L retrieval +λ2L generation ; Among them, L joint represents the total loss, L retrieval represents the first loss, L generation represents the second loss, λ1 represents the first loss weight, and λ2 represents the second loss weight.
13. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 12 are implemented.
14. A computer device comprising a memory, a processor and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 12 are implemented.
15. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 12 are implemented.
Citation Information
Patent Citations
Multi-modal fusion method based on cross attention gating unit
CN119046863A