A long document question-answering generation method and system based on large language models
By building HNSW index and position index methods, the calculation and memory usage of large language models are optimized, and the inefficiency problem of traditional Transformer models in long sequence processing is solved, and efficient long document question-and-answer generation is achieved.
Patent Information
- Application Number
- CN202411401128.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-10-09
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2044-10-09
AI Technical Summary
The traditional Transformer model faces computational efficiency and memory usage problems when processing long sequences, resulting in inefficiency in handling long documents and long conversations.
Using a question-and-answer generation method based on a large language model, an approximate nearest neighbor search is performed by constructing HNSW index and position index, and combining context window expansion and standard scaling dot product attention mechanism, optimize computing and memory usage.
It improves the computing efficiency and memory usage efficiency of long-sequence processing, ensures local coherence and global vision, avoids information fragmentation, and is suitable for processing large-scale data.
Smart Images

Figure CN119227810B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of artificial intelligence and natural language processing, and particularly to a method and system for generating long document question answers based on a large language model. Background Art
[0002] With the increasing complexity of natural language processing tasks, the need to process long documents and long conversations has become increasingly prominent. However, traditional Transformer models face serious computational efficiency and memory usage problems when dealing with long sequences. Summary of the Invention
[0003] To solve the technical problems existing in the above background art, the present invention provides a method for generating long document question answers based on a large language model, which helps to avoid the serious computational efficiency and memory usage problems faced by traditional Transformer models when dealing with long sequences.
[0004] To implement the above technical solution, in the first aspect, the present invention provides a method for generating long document question answers based on a large language model, including the following steps:
[0005] Step 1: Construct a prompt engineering template;
[0006] Step 2: Obtain a user question, insert the user question into the constructed prompt engineering template, and convert the prompt engineering template with the inserted user question into a question token sequence;
[0007] Step 3: Record the length QLEN of the question token sequence;
[0008] Step 4: Embed the question token sequence to obtain an embedded vector sequence E;
[0009] Step 5: Based on the obtained embedded vector sequence E, perform forward propagation calculations layer by layer in the large language model, and perform normalization on the calculation results of the last layer to generate a normalized layer output vector sequence;
[0010] Step 6: Based on the obtained normalized layer output vector sequence, generate a probability distribution of tokens, and select a token according to the token probability distribution to add to the question token sequence;
[0011] Step 7: Determine whether at least one of the following stop conditions is met. If so, collect all the tokens to form an answer text. If not, repeat steps 2 to 7;
[0012] Step 8: Process the formed answer text to generate a final answer.
[0013] Further, before the first step, it also includes: pre - constructing an HNSW index and a position index.
[0014] Further, the fifth step includes:
[0015] C1: Input the obtained embedding vector sequence E into the first layer of the large - language model, perform forward - propagation calculation in the first layer to obtain a layer output vector sequence OUT1, and input it into the second layer of the large - language model;
[0016] C2: For the m - th layer, where m≥2 and m is a positive integer; then, based on the vector sequence output from the (m - 1) - th layer of the large - language model, perform forward - propagation calculation to obtain a layer output vector sequence OUT m ;
[0017] C3: Determine whether the m - th layer of the large - language model is the last layer of the large - language model. If so, normalize the layer output vector sequence OUT m to obtain a normalized layer output vector sequence OUT. If not, use the layer output vector sequence OUT m as the input of the (m + 1) - th layer and continue to perform forward - propagation calculation in the (m + 1) - th layer.
[0018] Further, the pre - constructing the HNSW index and the position index includes:
[0019] Receive a long document input by the user;
[0020] Construct a prompt engineering template, input the content of the long document input by the user into the created prompt engineering, convert the prompt engineering with the inserted long - document content into a document token sequence, and record the length LEN of the document token sequence;
[0021] Based on the converted document token sequence, perform attention calculation layer by layer in the large - language model to create a triple for each token in the document token sequence, where each token's triple is (k i , v i , i), where i represents the i - th token in the document token sequence; k i represents the Key vector corresponding to the i - th token, and v i represents the Value vector corresponding to the i - th token;
[0022] Create a triple storage list for each layer of the large - language model;
[0023] Based on the triple storage list of each layer of the large - language model, construct an HNSW index and a position index.
[0024] Further, the step C2 includes:
[0025] Receiving the output vector sequence OUT of the upper layer m-1 , and using the layer normalization function LN to normalize the output vector sequence OUT m-1 to obtain a normalized vector sequence E m , where m E m-1 = LN(OUT
[0026] Performing self-attention calculation on the normalized vector sequence E m and linearly transforming the calculated self-attention to obtain a linearly transformed layer output vector sequence OUT m .
[0027] Further, the performing self-attention calculation on the normalized vector sequence E m and linearly transforming the calculated self-attention to obtain a linearly transformed layer output vector sequence OUT m includes:
[0028] Performing a linear transformation on the normalized vector sequence E m to generate a query vector matrix Q m , a key vector matrix K m and a value vector matrix V m ;
[0029] Using the query vector in the last column of the query vector matrix Q m to retrieve the top-N most similar key vectors from the pre-constructed HNSW index structure of the m-th layer;
[0030] Obtaining the positions of the tokens corresponding to the above top-N key vectors;
[0031] Using the position index structure to obtain all value vectors and key vectors within the range of the maximum number of neighbors M of the token position corresponding to each key vector;
[0032] Concatenating all the retrieved key vectors together in the order of the positions of the tokens corresponding to each key vector in the token sequence to obtain the matrix K m _hist, and concatenating all the retrieved value vectors together to obtain the matrix V m _hist;
[0033] Concatenating the obtained matrix K m _hist with the key vector matrix K m to obtain the matrix K' m, and the obtained matrix V m _hist and the value vector matrix V m are concatenated to obtain V m ';
[0034] Use rotary position encoding to add position information to the query vector matrix Q m and the matrix K' m to obtain the query vector matrix Q' m and the matrix K' m ;
[0035] Based on the matrix V' m and the query vector matrix Q' m with added position information and the matrix K' m , perform attention calculation and perform a linear transformation on the attention calculation result to obtain the linear transformation vector sequence O m ;
[0036] Based on the obtained linear transformation vector sequence O m and the normalized vector sequence E m , perform a residual connection to obtain the residual vector sequence A m ;
[0037] Normalize the obtained residual vector sequence A m to obtain the normalized residual vector sequence X m ;
[0038] Perform a two-layer feed-forward neural network linear calculation on the normalized residual vector sequence X m through the following formula;
[0039]
[0040] Add the output of the feed-forward neural network to the normalized residual vector sequence A m to obtain the layer output vector sequence OUT m .
[0041] In a second aspect, the present invention provides a long document question answering generation system based on a large language model, including:
[0042] A human-computer interaction module for receiving the long document and the user question input by the user, and displaying the final answer;
[0043] A sequence conversion module for converting the prompt engineering module inserted with the user question and the document content into a question token sequence and a document token sequence, and recording the lengths of the question token sequence and the document token sequence;
[0044] A question embedding module for embedding a question token sequence to obtain an embedded vector sequence E;
[0045] A propagation calculation module for performing forward propagation calculation layer by layer in a large language model based on the obtained embedded vector sequence E, and performing normalization on the calculation result of the last layer to generate a normalized layer output vector sequence;
[0046] A judgment module for judging whether a stop condition is satisfied;
[0047] An answer text generation module for collecting all tokens to form a complete answer text;
[0048] An answer text processing module for processing the formed answer text to generate a final answer;
[0049] A prompt engineering template construction module for constructing a prompt engineering template.
[0050] Furthermore, the system further includes:
[0051] An HNSW index and position index construction module for constructing an HNSW index and a position index.
[0052] In a third aspect, the present invention provides a computer-readable storage medium. The computer-readable storage medium includes a stored program, wherein when the program runs, it controls the device where the computer-readable storage medium is located to execute the above-mentioned question and answer generation method based on a large language model.
[0053] The beneficial effects of the present invention are as follows:
[0054] (1) The present invention constructs an independent vector database for each layer of the large language model to store (k, v, i) triples. Using the HNSW (Hierarchical Navigable Small World) algorithm, an efficient approximate nearest neighbor search structure is constructed with the k vector as the index. At the same time, a position-based digital index is created to achieve fast position retrieval, enabling faster search speed, strong scalability, suitable for processing large-scale data, and capable of completing approximate nearest neighbor search within an O(log n) time complexity, greatly improving the efficiency of subsequent retrieval.
[0055] (2) By performing approximate nearest neighbor search in the HNSW index, the most relevant N triples are retrieved, enabling the retrieval of the most relevant historical information for the current token in constant time, rather than linearly scanning the entire sequence. For each position in the retrieved triples, all (k, v) pairs within a fixed-size window around it (such as [-M, +M]) are obtained. Through context window expansion, local coherence is ensured, and the information fragmentation that may occur by relying solely on single-point retrieval is avoided. By concatenating the retrieved and expanded (k, v) pairs with the K and V matrices of the current input sequence to form new K' and V' matrices, using the standard scaled dot-product attention mechanism to calculate the attention, and only retaining the vectors corresponding to the input question sequence, both the global view is guaranteed and unnecessary calculations are avoided, thus solving the serious computational efficiency and memory usage problems faced when processing long sequences. Description of the Drawings
[0056] The specification drawings forming a part of the present invention are used to provide a further understanding of the present invention. The schematic embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation of the present invention.
[0057] Figure 1 It is a flowchart of the preparation stage of a question-answering generation method based on a large language model of the present invention.
[0058] Figure 2 It is a flowchart of the processing stage of a question-answering generation method based on a large language model of the present invention. Detailed Embodiments
[0059] The present invention will be further described below in conjunction with the drawings and embodiments.
[0060] It should be noted that the following detailed description is illustrative and is intended to provide a further description of the present invention. Unless otherwise specified, each technical and scientific term used in this embodiment has the same meaning as commonly understood by those of ordinary skill in the technical field to which the present invention belongs.
[0061] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to the present invention. As used herein, unless the context clearly indicates otherwise, the singular form is also intended to include the plural form. In addition, it should be understood that when the terms "comprising" and / or "including" are used in this specification, they indicate the presence of features, steps, operations, devices, components, and / or combinations thereof.
[0062] In the present invention, terms such as "upper", "lower", "left", "right", "front", "rear", "vertical", "horizontal", "side", "bottom", etc. indicate the orientation or positional relationship based on the orientation or positional relationship shown in the drawings. They are only relational terms determined for the convenience of describing the structural relationships of various components or elements of the present invention, and do not specifically refer to any component or element in the present invention. It should not be construed as a limitation to the present invention.
[0063] In the present invention, terms such as "fixed connection", "connected", "connected to" should be understood in a broad sense, which may mean a fixed connection, an integral connection or a detachable connection; it may be directly connected or indirectly connected through an intermediate medium. For relevant scientific research or technical personnel in this field, the specific meanings of the above terms in the present invention can be determined according to specific circumstances, and it should not be construed as a limitation to the present invention.
[0064] Embodiment 1:
[0065] This embodiment provides a long document question-answering generation method based on a large language model, including the following steps:
[0066] (1) Preparation stage: As Figure 1 shown, it includes the following steps:
[0067] S1: Receive the long document input by the user. Among them, the long document may include multiple texts.
[0068] S2: Construct a prompt engineering template, input the content of the long document input by the user into the created prompt engineering, convert the prompt engineering with the inserted long document content into a document token sequence, and record the length LEN of the document token sequence.
[0069] Specifically, the created prompt engineering (prompt) template is as follows:
[0070] The following is a long document. Please read it carefully and be prepared to answer questions:
[0071] {Document content}
[0072] Based on the above document, please answer the following questions.
[0073] S3: Based on the converted document token sequence, perform row attention calculation layer by layer in the large language model to obtain the Key vector and Value vector corresponding to each token in the document token sequence at each layer of the large language model; and based on the position of each token in the document token sequence, create (k i , v i , i) triples, where i represents the i-th token in the document token sequence; k iRepresents the Key vector corresponding to the i-th token, v i Represents the Value vector corresponding to the i-th token.
[0074] Among them, the position of each token in the document token sequence, and the sorting of the actual token in the document token sequence, is the ordinal number of the token.
[0075] S4: Create a triple storage list for each layer of the large language model to store the created triples.
[0076] Among them, each triple storage list includes LEN triples.
[0077] It should be noted that if the large language model has L layers, L triple storage lists can be obtained finally.
[0078] S5: Build an HNSW index based on the triple storage list of each layer of the large language model.
[0079] Specifically, it includes the following steps:
[0080] S51: Build an HNSW index structure based on the triple storage list of the large language model, and set the parameters of this HNSW index structure. The parameters include: the maximum number of neighbors M and the number of candidates efConstruction searched during construction.
[0081] S52: Obtain the triple storage list of each layer of the large language model, and add all the Key vectors in the triple storage list, as well as the Value vector information and position information in the token sequence corresponding to each Key vector, to the HNSW index structure to generate an HNSW index.
[0082] S53: Save the generated HNSW index to the database.
[0083] Among them, it should be noted that HNSW is a graphical index structure that organizes the high-dimensional vector space into a multi-layer navigable small-world graph. For each Key vector, calculate its distance from other Key vectors and establish connections in the graph.
[0084] S6: Build a position index based on the triple storage list of each layer of the large language model.
[0085] Specifically, it includes the following steps:
[0086] S61: Create a dictionary or data structure based on the triple storage list of the large language model.
[0087] S62: Map each token position in the triple storage list of the large language model to its corresponding Key vector and Value vector to generate a position index.
[0088] S63: Save the generated position index into the database.
[0089] (2) Processing stage, as Figure 2 shown, includes the following steps:
[0090] S7: Obtain the user's question, insert the user's question into the constructed prompt engineering template, and convert the prompt engineering template with the inserted user's question into a question token sequence.
[0091] Specifically, the created user question prompt engineering (prompt) template is as follows:
[0092] {user question}
[0093] Specifically, for example, use the tokenizer of the large language model (such as the Llama 2 model) to convert the prompt engineering template with the inserted user's question into a question token sequence.
[0094] S8: Record the length QLEN of the question token sequence.
[0095] Here, it should be noted that the length of the token sequence is essentially the number of tokens in the token sequence.
[0096] S9: Embed the question token sequence to obtain an embedded vector sequence E.
[0097] Here, the shape of the embedded vector sequence E is [QLEN, d model , d model is the dimension of the embedded vector sequence, which is a parameter set in the large language model configuration file.
[0098] S10: Based on the obtained embedded vector sequence E, perform forward propagation calculations layer by layer in the large language model, and perform normalization on the calculation results of the last layer to generate a normalized layer output vector sequence.
[0099] Here, it should be noted that different large language models have different numbers of layers. For example, the Llama2-7B large language model has 32 layers, the Llama2-13B large language model has 40 layers, and the Qwen2.5-3B large language model has 36 layers.
[0100] In this embodiment, the large language model uses the Llama2 large language model. Assuming that the large language model has L layers as an example; L is an integer greater than or equal to 2. Forward propagation calculations are performed layer by layer in the large language model, specifically including the following steps:
[0101] S10-1: First, input the obtained embedding vector sequence E into the first layer of the large language model and perform the following operations in the first layer:
[0102] A1: Use the layer normalization function LN to normalize the embedding vector sequence E to obtain a normalized vector sequence E1, where E1 = LN(E).
[0103] A2: Perform self-attention calculation on the normalized vector sequence E1 and perform a linear transformation on the calculated self-attention to obtain the linearly transformed self-attention.
[0104] Specifically, it includes the following steps:
[0105] A2-1: Perform a linear transformation on the normalized vector sequence E1 to generate a query vector matrix Q1, a key vector matrix K1, and a value vector matrix V1 for the user question.
[0106] Among them, Q1 = Linearq(E1); K1 = Lineark(E1); V1 = Linearv(E1); the shapes of Q1, K1, and V1 are all [QLEN, d k , where d k is a pre-set dimension.
[0107] A2-2: Use the query vector in the last column of the query vector matrix Q1 to retrieve the top-N most similar key vectors in the pre-constructed HNSW index structure of the first layer.
[0108] A2-3: Obtain the positions of the tokens corresponding to the above top-N key vectors.
[0109] A2-4: Use the position index structure to obtain all value vectors and key vectors within the range of the maximum number of neighbors M of the token position corresponding to each key vector.
[0110] Specifically, assuming that the token position corresponding to one of the key vectors is i, then obtain all value vectors and key vectors within the range of [i - M, i + M].
[0111] Among them, the value of M is pre-set according to actual needs.
[0112] Among them, N is a predefined parameter and can be set according to actual needs.
[0113] A2-5: Concatenate all retrieved Key vectors together in the order of the positions of the tokens corresponding to each Key vector in the token sequence to obtain matrix K1_hist (with a shape of [L_hist, d k ), and concatenate all retrieved Value vectors together to obtain matrix V1_hist (also with a shape of [L_hist, d k ).
[0114] A2-6: Concatenate the obtained matrix K1_hist with the key vector matrix K1 to obtain matrix K'1, and concatenate the obtained matrix V1_hist with the value vector matrix V1 to obtain V'1.
[0115] Specifically, concatenate the obtained matrix K1_hist in front of the key vector matrix K1 to obtain matrix K'1, and concatenate matrix V1_hist in front of the value vector matrix V1 to obtain matrix V'1.
[0116] Among them, the shapes of matrices K'1 and V'1 are both [QLEN + L_hist, d k .
[0117] A2-7: Use rotational position encoding (e.g., PoPE) to add position information to the query vector matrix Q1 and matrix K'1 to obtain the query vector matrix Q'1 and matrix K'1.
[0118] Specifically, adding rotational position encoding to the query vector matrix Q1 and matrix K'1 can ensure that when calculating self-attention, the relative position information between tokens is included, thus achieving a more comprehensive and complete semantic understanding.
[0119] A2-8: Based on matrix V'1 and the query vector matrix Q'1 and matrix K'1 with added position information, perform attention calculation and perform a linear transformation on the attention calculation result to obtain the linear transformation vector sequence O1.
[0120] Among them, the attention calculation formula is as follows:
[0121]
[0122] The formula for performing a linear transformation on the attention calculation result to obtain the linear transformation vector sequence O1 is as follows:
[0123]
[0124] A2-9: Based on the obtained linear transformation vector sequence O1 and the normalized vector sequence E1, perform a residual connection to obtain the residual vector sequence A1;
[0125] Among them, the residual connection formula is as follows:
[0126] A1 = E1 + O1.
[0127] A2 - 10: Normalize the obtained residual vector sequence A1 to obtain the normalized residual vector sequence X1.
[0128] Among them, the formula for normalizing the obtained residual vector sequence A1 is as follows:
[0129] X1 = LN(A1).
[0130] A2 - 11: Perform two - layer feed - forward neural network linear calculation on the normalized residual vector sequence X1 through the following formula;
[0131]
[0132] A2 - 12: Add the output of the feed - forward neural network to the normalized residual vector sequence A1 to obtain the layer output vector sequence OUT1, and use the layer output vector sequence OUT1 as the input of the second layer.
[0133] Specifically, OUT1 = FFN(X1) + A1.
[0134] S10 - 2: For the m - th layer, where L≥m≥2 and m is a positive integer, the following operations are performed:
[0135] B1: Receive the output vector sequence OUT of the previous layer m-1 , and use the layer normalization function LN to normalize the output vector sequence OUT m-1 to obtain the normalized vector sequence E m , where E m = LN(OUT m-1 ).
[0136] B2: Perform self - attention calculation on the normalized vector sequence E m and linearly transform the calculated self - attention to obtain the linearly transformed layer output vector sequence OUT m .
[0137] Specifically, it includes the following steps:
[0138] B2 - 1: Linearly transform the normalized vector sequence E m to generate the query vector matrix Q m of the user's question, the key vector matrix K m and the value vector matrix V m .
[0139] Among them, Q m = Linearq (E m );K m = Linear k (E m );V m = Linear v (E m );Q m 、K m 、V m are all in the shape of [QLEN, d k , where d k is a dimension set in advance.
[0140] B2-2: Retrieve the top-N most similar Key vectors from the pre-constructed HNSW index structure of the m-th layer using the query vector in the last column of the query vector matrix Q m .
[0141] B2-3: Obtain the positions of the tokens corresponding to the above top-N Key vectors.
[0142] B2-4: Use the position index structure to obtain all Value vectors and Key vectors within the range of the maximum number of neighbors M corresponding to each Key vector.
[0143] Specifically, assuming that the token position corresponding to one of the Key vectors is i, then obtain all Value vectors and Key vectors within the range of positions [i - M, i + M].
[0144] Among them, the value of M is set in advance according to actual needs.
[0145] Among them, N is a predefined parameter, which can be set according to actual needs.
[0146] B2-5: Concatenate all the retrieved Key vectors together in the order of the positions of the tokens corresponding to each Key vector in the token sequence to obtain the matrix K m _hist (whose shape is [L_hist, d k ), and concatenate all the retrieved Value vectors together to obtain the matrix V m _hist (also in the shape of [L_hist, d k ).
[0147] B2-6: Concatenate the obtained matrix K m _hist with the key vector matrix K m to obtain the matrix K' m , and concatenate the obtained matrix V m_hist and the value vector matrix V m Concatenate them to obtain V m '.
[0148] Specifically, concatenate the obtained matrix K m _hist to the key vector matrix K m to obtain the matrix K' m , and concatenate the matrix V m _hist to the value vector matrix V m to obtain the matrix V' m .
[0149] Among them, the shapes of the matrices K' m and V' m are both [QLEN + L_hist, d k .
[0150] B2-7: Use rotational position encoding (e.g., PoPE) to add position information to the query vector matrix Q m and the matrix K' m to obtain the query vector matrix Q' m and the matrix K' m .
[0151] Specifically, adding rotational position encoding to the query vector matrix Q m and the matrix K' m can ensure that when calculating self-attention, the relative position information between tokens is included, thus achieving a more comprehensive and complete semantic understanding.
[0152] B2-8: Based on the matrix V' m and the query vector matrix Q' m with added position information and the matrix K' m , perform attention calculation and perform a linear transformation on the attention calculation result to obtain the linear transformation vector sequence O m .
[0153] Among them, the attention calculation formula is as follows:
[0154]
[0155] The formula for performing a linear transformation on the attention calculation result to obtain the linear transformation vector sequence O m is as follows:
[0156]
[0157] B2-9: Based on the obtained linear transformation vector sequence O m and the normalized vector sequence E m, perform residual connection to obtain a sequence of residual vectors A m ;
[0158] Among them, the residual connection formula is as follows:
[0159] A m = E m + O m .
[0160] B2-10: Normalize the obtained sequence of residual vectors A m to obtain a sequence of normalized residual vectors X m .
[0161] Among them, the formula for normalizing the obtained sequence of residual vectors A is as follows:
[0162] X m = LN(A m ).
[0163] B2-11: Perform two-layer feedforward neural network linear calculation on the sequence of normalized residual vectors X m ;
[0164]
[0165] B2-12: Add the output of the feedforward neural network to the sequence of normalized residual vectors A m to obtain a sequence of layer output vectors OUT m .
[0166] Specifically, OUT m = FFN(X m ) + A m .
[0167] B3: Determine whether layer m is the last layer of the large language model. If so, normalize the sequence of layer output vectors OUT m to obtain a sequence of normalized layer output vectors OUT. If not, use the sequence of layer output vectors OUT m as the input of layer m+1, and execute steps B1 to B2 in layer m+1.
[0168] S11: Based on the obtained sequence of normalized layer output vectors OUT, generate a probability distribution of tokens, select a token according to the token probability distribution, and add the selected token to the sequence of problem tokens.
[0169] Specifically, the standardized layer output vector sequence OUT is passed to the large language model head to generate the probability distribution result of the next token. Based on the generated probability distribution result of the next token, a sampling strategy (such as greedy decoding or beam search) selects the next token and adds the selected token to the question token sequence.
[0170] S12: Determine whether at least one of the following stop conditions is met. If so, collect all tokens to form the complete answer text; if not, repeat steps S8 to S12.
[0171] Among them, the stop conditions include: whether the end symbol is generated and whether the question token sequence reaches the predefined token sequence length.
[0172] S13: Process the formed answer text to generate the final answer.
[0173] Among them, processing the formed answer text includes: removing special tokens and formatting, etc.
[0174] Embodiment 2:
[0175] This embodiment provides a long document question-answering generation system based on a large language model, including:
[0176] (1) A human-computer interaction module, which is used to receive the long document and user question input by the user, and display the final answer.
[0177] (2) An HNSW index and position index construction module, which is used to construct the HNSW index and position index of the large language model.
[0178] The HNSW index and position index construction module includes:
[0179] A prompt engineering construction module, which is used to construct a prompt engineering template.
[0180] An engineering conversion unit, which is used to convert the prompt engineering template into a document token sequence and record the length of the document token sequence.
[0181] A triple creation unit, which is used to perform attention calculations layer by layer in the large language model based on the converted document, so as to obtain the Key vector and Value vector corresponding to each token in the document token sequence in each layer of the large language model; and create (k i , v i , i) triples based on the position of each token in the document token sequence.
[0182] Storage list creation module, which is used to create a triple storage list for each layer of the large language model for storing the created (k i , v i , i) triples.
[0183] HNSW index structure construction unit, which is used to construct the HNSW index structure;
[0184] HNSW index structure parameter setting unit, which is used to set the parameters of the HNSW index structure.
[0185] HNSW index construction unit, which obtains the triple storage list of each layer of the large language model, and adds all the Key vectors in the triple storage list, as well as the Value vector information corresponding to each Key vector and the position information in the token sequence, to the HNSW index structure to generate the HNSW index.
[0186] Position index structure construction unit, which is used to create a dictionary or data structure.
[0187] Position index creation unit, which is used to map each token position in the triple storage list of each layer of the large language model to its corresponding Key vector and Value vector to generate the position index.
[0188] Database, which is used to store the position index and the HNSW index.
[0189] (III) Sequence conversion module, which is used to convert the prompt engineering module inserted with the user question or document content into a question token sequence and record the length of the question token sequence;
[0190] (IV) Question embedding module, which is used to embed the question token sequence to obtain the embedded vector sequence E;
[0191] (V) Propagation calculation module, which is used to perform forward propagation calculation layer by layer in the large language model based on the obtained embedded vector sequence E, and perform normalization on the calculation result of the last layer to generate the normalized layer output vector sequence.
[0192] (VI) Sequence selection module, which is used to generate the probability distribution of tokens based on the obtained normalized layer output vector sequence, select a token according to the token probability distribution, and add the selected token to the question token sequence.
[0193] (VII) Judgment module, which is used to judge whether the stop condition is met;
[0194] (VIII) Answer text generation module, which is used to collect all the tokens to form the complete answer text;
[0195] (9) Answer text processing module, configured to process the composed answer text to generate a final answer.
[0196] (10) Prompt engineering template construction module, configured to construct a prompt engineering template.
[0197] Embodiment 3:
[0198] This embodiment provides a computer-readable storage medium, which includes a stored program. When the program runs, it controls the device where the computer-readable storage medium is located to execute the method for generating long document Q&A based on a large language model described in Embodiment 2.
[0199] For the same or similar parts among the various embodiments in this specification, reference can be made to each other. In particular, for the terminal embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and reference can be made to the descriptions in the method embodiments for the relevant parts.
[0200] In several embodiments provided by the present invention, it should be understood that the disclosed system and method can
[0201] be implemented in other ways. For example, the system embodiments described above are only illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some interfaces. The indirect coupling or communication connection of the system or unit can be in an electrical, mechanical or other form.
[0202] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place, or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0203] In addition, it should be noted that the flowcharts in the accompanying drawings illustrate the methods of the embodiments of the present disclosure. In the corresponding descriptions in the flowcharts or block diagrams in the accompanying drawings, the operations or steps corresponding to different blocks may also occur in an order different from that disclosed in the description. Sometimes, there is no specific order between different operations or steps. For example, two consecutive operations or steps can actually be executed substantially in parallel, and sometimes they can also be executed in the reverse order, which may depend on the functions involved. Each block in the block diagram and / or flowchart, as well as the combination of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system that performs the specified functions or actions, or can be implemented by a combination of dedicated hardware and computer instructions.
[0204] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. For those skilled in the art, the present invention can have various changes and modifications. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. A long document question-answering generation method based on a large language model, characterized in that, It includes the following steps: Step 1: Construct a prompt engineering template; Step 2: Obtain the user's question, insert the user's question into the constructed prompt engineering template, and convert the prompt engineering template with the inserted user's question into a question token sequence; Step 3: Record the length QLEN of the question token sequence; Step 4: Embed the question token sequence to obtain an embedded vector sequence E; Step 5: Based on the obtained embedded vector sequence E, perform forward propagation calculations layer by layer in the large language model, and perform normalization on the calculation results of the last layer to generate a normalized layer output vector sequence; Step 6: Based on the obtained normalized layer output vector sequence, generate a probability distribution of tokens, and select a token according to the token probability distribution to add to the question token sequence; Step 7: Determine whether at least one of the following stop conditions is satisfied. If so, collect all the tokens to form an answer text. If not, repeat Steps 2 to 7; Step 8: Process the formed answer text to generate a final answer; Before the above Step 1, it also includes: pre-constructing an HNSW index and a position index; The pre-constructing of the HNSW index and the position index includes: Receiving a long document input by the user; Constructing a prompt engineering template, inputting the content of the long document input by the user into the created prompt engineering, converting the prompt engineering with the inserted long document content into a document token sequence, and recording the length LEN of the document token sequence; Based on the converted document token sequence, perform attention calculations layer by layer in the large language model to create triples for each token in the document token sequence, where each token triple is (k i , v i , i), where i represents the i-th token in the document token sequence; k i represents the Key vector corresponding to the i-th token, and v i represents the Value vector corresponding to the i-th token; Creating a triple storage list for each layer of the large language model; Based on the triple storage list of each layer of the large language model, constructing an HNSW index and a position index; The above Step 5 includes: C1: Input the obtained embedded vector sequence E into the first layer of the large language model, perform forward propagation calculations in the first layer to obtain a layer output vector sequence OUT1 and input it into the second layer of the large language model; C2: For the m-th layer, where m ≥ 2 and m is a positive integer; then, based on the vector sequence output by the (m - 1)-th layer of the large language model, perform forward propagation calculations to obtain the layer output vector sequence OUT m ; The above C2 includes: Receive the output vector sequence OUT from the upper layer m-1 , and use the layer normalization function LN to normalize the output vector sequence OUT m-1 to obtain the normalized vector sequence E m , where E m =LN(OUT m-1 ); For the standardized vector sequence E m Perform self-attention calculation and linearly transform the calculated self-attention to obtain the linearly transformed layer output vector sequence OUT m ; The standardized vector sequence E m Performs self-attention calculation and linearly transforms the calculated self-attention to obtain a linearly transformed layer output vector sequence OUT m Includes: Perform a linear transformation on the standardized vector sequence E m to generate a query vector matrix Q m , a key vector matrix K m and a value vector matrix V m ; Among them, Q m = Linear q (E m );K m = Linear k (E m );V m = Linear v (E m );Q m 、K m 、V m are all in the shape of [QLEN, d k , where d k is a preset dimension; Use the query vector in the last column of the query vector matrix Q m to retrieve the top-N most similar key vectors from the pre-constructed HNSW index structure of the m-th layer; Obtaining the positions of the tokens corresponding to the above TOP -N Key vectors; Using the position index structure, obtaining all the Value vectors and Key vectors within the range of the maximum number of neighbors M of the token positions corresponding to each Key vector; Concatenate all the retrieved Key vectors together in the order of the positions of the tokens corresponding to each Key vector in the token sequence to obtain matrix K m _hist (with shape [L_hist, d k ), and concatenate all the retrieved Value vectors together to obtain matrix V m _hist (also with shape [L_hist, d k ); The obtained matrix K m _hist and the key vector matrix K m are concatenated to obtain matrix K' m , and the obtained matrix V m _hist and the value vector matrix V m are concatenated to obtain V' m ; Use rotational position encoding (e.g., PoPE) for the query vector matrix Q m and matrix K' m Add position information to obtain the query vector matrix Q' m and matrix K' m; Based on matrix V' m and the query vector matrix Q' with added position information m and matrix K' m , perform attention calculation and perform a linear transformation on the attention calculation result to obtain the linear transformation vector sequence O m ; Based on the obtained sequence of linear transformation vectors O m and the sequence of normalized vectors E m , perform residual connection to obtain the residual vector sequence A m ; The obtained residual vector sequence A m is normalized to obtain a normalized residual vector sequence X m ; Perform a two-layer feedforward neural network linear calculation on the standardized residual vector sequence X m ; Add the output of the feedforward neural network to the standardized residual vector sequence A m to obtain the layer output vector sequence OUT m .
2. The long document question-answering generation method based on a large language model according to claim 1, wherein The above Step 5 also includes: C3: Determine whether the m-th layer of the large language model is the last layer of the large language model. If so, normalize the layer output vector sequence OUT m to obtain the normalized layer output vector sequence OUT. If not, use the layer output vector sequence OUT m as the input for the (m+1)-th layer and continue to perform forward propagation calculations in the (m+1)-th layer.
3. A long document Q&A generation system based on a large language model, based on the long document Q&A generation method based on a large language model according to any one of claims 1-2, characterized in that, including: A human-computer interaction module, used to receive the long document and the user's question input by the user, and display the final answer; A sequence conversion module, used to convert the prompt engineering module with the inserted user's question and document content into a question token sequence and a document token sequence, and record the length of the question token sequence and the length of the document token sequence; A question embedding module, used to embed the question token sequence to obtain an embedded vector sequence E; A propagation calculation module, used to perform forward propagation calculations layer by layer in the large language model based on the obtained embedded vector sequence E, and perform normalization on the calculation results of the last layer to generate a normalized layer output vector sequence; A judgment module, used to judge whether the stop condition is satisfied; An answer text generation module, used to collect all the tokens to form a complete answer text; Answer text processing module, used to process the composed answer text to generate the final answer; Prompt engineering template construction module, used to construct a prompt engineering template; HNSW index and position index construction module, used to construct an HNSW index and a position index; The pre-constructed HNSW index and position index include: Receive the long document input by the user; Construct a prompt engineering template, input the content of the long document received by the user into the created prompt engineering, convert the prompt engineering with the long document content inserted into a document token sequence, and record the length LEN of the document token sequence; Based on the converted document token sequence, perform attention calculations layer by layer in the large language model to create triples for each token in the document token sequence, where each triple for a token is (k i , v i , i), where i represents the i-th token in the document token sequence; k i represents the Key vector corresponding to the i-th token, and v i represents the Value vector corresponding to the i-th token; Create a triple storage list for each layer of the large language model; Based on the triple storage list of each layer of the large language model, construct an HNSW index and a position index; The forward propagation calculation is performed layer by layer in the large language model based on the obtained embedding vector sequence E, and the calculation result of the last layer is normalized to generate a normalized layer output vector sequence, including: C1: Input the obtained embedding vector sequence E into the first layer of the large language model, perform forward propagation calculation in the first layer to obtain a layer output vector sequence OUT1 and input it into the second layer of the large language model; C2: For the m-th layer, where m ≥ 2 and m is a positive integer; then, based on the vector sequence output by the (m - 1)-th layer of the large language model, perform forward propagation calculations to obtain the layer output vector sequence OUT m ; The C2 includes: Receive the output vector sequence OUT from the upper layer m-1 , and use the layer normalization function LN to normalize the output vector sequence OUT m-1 to obtain the normalized vector sequence E m , where E m =LN(OUT m-1 ); For the standardized vector sequence E m Perform self-attention calculation and linearly transform the calculated self-attention to obtain the linearly transformed layer output vector sequence OUT m ; The standardized vector sequence E m performs self-attention calculation and linearly transforms the calculated self-attention to obtain a linearly transformed layer output vector sequence OUT m including: Perform a linear transformation on the standardized vector sequence E m to generate the query vector matrix Q m , key vector matrix K m and value vector matrix V m ; Among them, Q m = Linear q (E m ); K m = Linear k (E m ); V m = Linear v (E m ); Q m , K m , V m are all in the shape of [QLEN, d k , where d k is a preset dimension; Use the query vector in the last column of the query vector matrix Q m to retrieve the top-N most similar key vectors from the pre-built HNSW index structure of the m-th layer; Obtain the positions of the tokens corresponding to the above TOP -N Key vectors; Use the position index structure to obtain all Value vectors and Key vectors within the range of the maximum number of neighbors M of the token positions corresponding to each Key vector; Concatenate all the retrieved Key vectors together in the order of the positions of the tokens corresponding to each Key vector in the token sequence to obtain the matrix K m _hist (with shape [L_hist, d k ), and concatenate all the retrieved Value vectors together to obtain the matrix V m _hist (also with shape [L_hist, d k ) The obtained matrix K m _hist and the key vector matrix K m are concatenated to obtain matrix K' m , and the obtained matrix V m _hist and the value vector matrix V m are concatenated to obtain V' m ; Use rotational positional encoding (e.g., PoPE) for the query vector matrix Q m and the matrix K' m Add positional information to obtain the query vector matrix Q' m and the matrix K' m; Based on matrix V' m and the query vector matrix Q' with added position information m and matrix K' m , perform attention calculation and perform a linear transformation on the attention calculation result to obtain a sequence of linear transformation vectors O m ; Based on the obtained sequence of linear transformation vectors O m and the sequence of normalized vectors E m , perform residual connection to obtain the sequence of residual vectors A m ; For the obtained residual vector sequence A m perform standardization to obtain the standardized residual vector sequence X m ; Perform a two - layer feed - forward neural network linear calculation on the standardized residual vector sequence X m ; Add the output of the feedforward neural network to the standardized residual vector sequence A m to obtain the layer output vector sequence OUT m .
4. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes a stored program, wherein when the program runs, it controls the device where the computer-readable storage medium is located to execute the large language model-based question and answer generation method according to any one of claims 1 to 2.
Citation Information
Patent Citations
Networking question and answer method based on large language model
CN118227890A
Intelligent legal question and answer method based on retrieval enhanced language model
CN118277538A