Training Method, Question-Answering Method, Device, Medium and Product of Large Language Model
By obtaining long text training data and increasing the rotation angle base of the rotation position code, the large language model is retrained, and the large language model answers incomplete and inaccurate questions when comparing long text and multiple documents, achieving a more efficient question-and-answer system.
Patent Information
- Application Number
- CN202411282695.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-13
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2044-09-13
AI Technical Summary
When existing large language models deal with long text and multi-document comparisons, the length of the input text sequence is limited, resulting in incomplete and accurate answers.
By obtaining long text training data and increasing the rotation angle base number of pre-trained large language models encoded by the rotation position of the pre-trained large language model, the large language model is retrained to expand the maximum length of its input text sequence.
It realizes the completeness and accuracy of answers of large language models on questions that rely on long text and multi-document comparisons, simplifies the system structure, and improves the effectiveness of the question-and-answer system.
Smart Images

Figure CN118798303B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and in particular, to a training method, a question-answering method, a device, a medium, and a product for a large language model. Background Art
[0002] The Retrieval-Augmented Generation (RAG) technology is a very widely used application solution for large language models. By retrieving relevant materials and then assembling them according to preset instructions, it is input into the large language model for understanding, and finally answers the user's questions. It is commonly used in various knowledge Q&A assistants. Compared with the direct answer of the large language model, the answer in the RAG method utilizes the knowledge retrieved, which can ensure the timeliness of the knowledge and reduce the problem of inaccurate results caused by the large language model fabricating by itself.
[0003] The RAG solution in the related technology splits the collected documents into multiple text chunks (chunks) according to a specific chunking method and stores them in databases such as a vector library and a search engine library (for example, Elasticsearch library). When a question raised by the user is obtained, the most relevant several text chunks are retrieved respectively through parallel queries of multiple libraries, and a sorting model is called to sort the text chunks recalled from multiple paths. The text chunks with the top rankings are input into the large language model, and the answer to the question is obtained based on the inference of the large language model. However, since the length of the text sequence input into the large language model is limited, the inference effect of questions that require summarization of long documents or comparison and summarization of multiple documents is not good. Summary of the Invention
[0004] Embodiments of this application provide a training method, a question-answering method, a device, a medium, and a product for a large language model to expand the maximum text sequence length input into the large language model and improve the answer integrity and accuracy of the large language model in the problems of long text dependence and multi-document comparison dependence.
[0005] In a first aspect, an embodiment of this application provides a training method for a large language model, including: obtaining long text training data, where the sequence length of the long text training data is greater than the maximum length of the input text sequence of the pre-trained large language model; increasing the rotation angle base number of the rotary position encoding of the pre-trained large language model to obtain a modified pre-trained large language model; using the long text training data to train the modified pre-trained large language model to obtain a trained large language model.
[0006] Second aspect, an embodiment of the present application provides a question-answering method based on a large language model, including: obtaining question information and a task instruction, and querying a plurality of text blocks related to the question information; splicing the question information, the task instruction, and the plurality of text blocks to obtain a spliced text; the sequence length of the spliced text is greater than the maximum length of the input text sequence of the pre-trained large language model; inputting the spliced text into the trained large language model to obtain reply information output by the trained large language model; wherein, the trained large language model is trained by using the training method in the embodiment of the present application.
[0007] Third aspect, an embodiment of the present application provides an electronic device, including a memory, a processor, and a computer program stored on the memory. When the processor executes the computer program, the method in any one of the above is implemented.
[0008] Fourth aspect, an embodiment of the present application provides a computer-readable storage medium, in which a computer program is stored. When the computer program is executed by a processor, the method in any one of the above is implemented.
[0009] Fifth aspect, an embodiment of the present application provides a computer program product, which includes a computer program. When the computer program is executed by a processor, the method in any one of the above is implemented.
[0010] Compared with the prior art, the present application has the following advantages:
[0011] The present application provides a training method, a question-answering method, a device, a medium, and a product of a large language model, obtaining long text training data, and the sequence length of the long text training data is greater than the maximum length of the input text sequence of the pre-trained large language model; increasing the rotation angle base of the rotary position encoding of the pre-trained large language model to obtain a modified pre-trained large language model; using the long text training data to train the modified pre-trained large language model to obtain a trained large language model. In this embodiment, by obtaining long text training data and increasing the rotation angle base of the rotary position encoding, the pre-trained large language model is trained. Since the rotary position encoding generates position encoding vectors through mathematical transformation, increasing the rotation angle base of the rotary position encoding can change the way of mathematical transformation, thereby affecting the generation of position encoding vectors, so that the trained large language model can still effectively capture position information when processing longer sequences, thus allowing the large language model to process longer input text sequences, achieving the amplification of the length of the input text sequence, and improving the answer integrity and accuracy of the large language model in the problems of long text dependence and multi-document comparison dependence.
[0012] The above description is only an overview of the technical solution of this application. In order to understand the technical means of this application more clearly, it can be implemented according to the content of the specification. And in order to make the above and other purposes, features and advantages of this application more obvious and understandable, the specific implementation manners of this application are specifically exemplified below. Description of the Drawings
[0013] In the drawings, unless otherwise specified, the same reference numerals throughout the several views denote the same or similar parts or elements. These drawings are not necessarily drawn to scale. It should be understood that these drawings only depict some embodiments according to this application and should not be regarded as limiting the scope of this application.
[0014] Figure 1 It is a schematic diagram of the splicing of input text blocks of the large language model in the related art.
[0015] Figure 2 It is a schematic diagram of the splicing of input text blocks of the large language model according to an embodiment of this application.
[0016] Figure 3 It is a flowchart of the training method of the large language model according to an embodiment of this application.
[0017] Figure 4 It is a flowchart of the training method of the large language model according to an embodiment of this application.
[0018] Figure 5 It is a flowchart of the question-answering method based on the large language model according to an embodiment of this application.
[0019] Figure 6 It is a flowchart of the question-answering method based on the large language model according to an embodiment of this application.
[0020] Figure 7 It is a structural block diagram of the training device of the large language model according to an embodiment of this application.
[0021] Figure 8 It is a structural block diagram of the question-answering device based on the large language model according to an embodiment of this application.
[0022] Figure 9 It is a block diagram of the electronic device for implementing the embodiments of this application. Detailed Description of the Invention
[0023] In the following, only some exemplary embodiments are briefly described. As those skilled in the art can recognize, the described embodiments can be modified in various different ways without departing from the concept or scope of this application. Therefore, the drawings and the description are considered to be exemplary in nature and not restrictive.
[0024] To facilitate the understanding of the technical solutions of the embodiments of the present application, the related technologies of the embodiments of the present application are described below. The following related technologies can be arbitrarily combined with the technical solutions of the embodiments of the present application as optional solutions, and they all fall within the protection scope of the embodiments of the present application.
[0025] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in the present application are all information and data authorized by the user or fully authorized by all parties. Moreover, the collection, use, and processing of relevant data need to comply with the relevant laws, regulations, and standards of relevant countries and regions, and corresponding operation entrances are provided for users to choose to authorize or refuse.
[0026] In the related technology, the RAG technology retrieves multiple text blocks related to the question input by the user through the retrieval component. Since the input text (also called input context) sequence length of the open-source large language model is limited, generally 8k - 32k, where k = 1024, that is, the length of the input text sequence is at most 32×1024 tokens. When the sequence length of multiple text blocks exceeds 32k, they cannot be input into the large language model. Therefore, the sorting model is used to sort multiple text blocks, and several text blocks with higher rankings are selected and concatenated with the task instruction and question information and then input into the large language model for processing.
[0027] Figure 1 It is a splicing schematic diagram of the input text blocks of the large language model in the related technology. Among them, E - text block 1, E - text block 2... E - text block N represent N text blocks related to the question information input by the user retrieved by the text keyword recall method. V - text block 1, V - text block 2... V - text block M represent M text blocks related to the question information input by the user retrieved by the vector recall method. The sorting model is used to sort the text blocks, and 3 text blocks with higher rankings, R - text block 1, R - text block 2, and R - text block 3, are selected. The selected text blocks are concatenated with the task instruction and the question information input by the user and then input into the large language model for processing to obtain the reply information. Among them, the task instruction (Instruction) is the guidance or command required to direct the large language model to execute a specific task, indicating the specific task the user wants the large language model to complete or the type of output required. For example, you are an xxx expert. Please answer the user's question xxx according to the retrieved xxx and in accordance with the xxx requirements.
[0028] Since the length of the text sequence input to the large language model is limited, in related technologies, only the top-ranked text blocks can be selected. For questions that require summarizing long documents or comparing and summarizing multiple documents to answer, the inference results of the large language model are not complete and accurate enough.
[0029] In the embodiments of the present application, first, long text training data is obtained. The sequence length of the long text training data is greater than the maximum length of the input text sequence of the pre-trained large language model. Specifically, training corpora with long contexts are collected, such as data from network datasets or book document data, and then cleaned to a preset length, such as 128k (k = 1024). Then, the rotation angle base of the rotary position encoding of the pre-trained large language model is increased to obtain a modified pre-trained large language model. Specifically, the rotation angle base of the rotary position encoding (i.e., the base value of the rotary position encoding) is modified. The rotation angle base is generally defaulted to 10,000 (ten thousand) and can be modified to 1,000,000 (one million). Since the rotary position encoding generates position encoding vectors through mathematical transformations, increasing the rotation angle base of the rotary position encoding can change the way of mathematical transformation, thereby affecting the generation of position encoding vectors, enabling the trained large language model to still effectively capture position information when processing longer sequences, and thus allowing the large language model to process longer input text sequences. The modified pre-trained large language model is trained using the long text training data to obtain a trained large language model, that is, a general large language model.
[0030] Then, the general large language model is fine-tuned. Long text training data in a preset domain is obtained, and the trained large language model is trained using the long text training data in the preset domain to obtain a large language model corresponding to the preset domain.
[0031] Finally, position expansion is performed. Specifically, in the configuration file of the large language model corresponding to the preset domain (i.e., the first large language model), the maximum position number of the large language model (the maximum position number represents the maximum text sequence length that the large language model can process) is increased to obtain a large language model with position expansion (i.e., the second large language model). For example, using the rotary position encoding expansion method YaRN (Yet another RoPE extensioN) method, the code of the model inference framework vllm, in the configuration file, if the maximum length of the large language model context (original_max_position_embeddings) is 32k, position expansion is performed by length extrapolation, and the extrapolation multiple (factor) is set to 8 times, that is, the maximum length of the input text sequence of the modified large language model is 32k * 8 = 256k.
[0032] Since the large language model in the embodiments of this application is trained using long text data, increases the rotation angle base of the rotary position encoding, and expands the position of the input text sequence length of the large language model, a sorting model is not required, and text blocks can be sorted according to the relevance score. Since the large language model is more sensitive to the text closer to the task instruction and the part closer to the question information, and since the text blocks obtained by the vector recall method are more accurate than the text obtained by the text keyword recall method, therefore, the text blocks E-Text Block 1, E-Text Block 2... E-Text Block N obtained by the text keyword recall method are sorted from high to low according to the relevance score with the question information, and the text blocks V-Text Block 1, V-Text Block 2... V-Text Block M obtained by the vector recall method are sorted from low to high according to the relevance score with the question information. The text blocks E-Text Block 1, E-Text Block 2... E-Text Block N are concatenated after the task instruction, and the text blocks V-Text Block 1, V-Text Block 2... V-Text Block M are concatenated before the question information. Figure 2 This is a splicing schematic diagram of the input text blocks of the large language model according to an embodiment of this application. In addition, in order to make the large language model more easily distinguish the content at different positions (task instruction, text block, question information), prefix tokens Pre-Tokens and suffix tokens Post-Tokens are respectively inserted as the first placeholder and the second placeholder. The number of placeholders for different large language models is different. For example, Pre-Tokens can be 400, and Post-Tokens can be 100. Finally, the spliced text is input to the large language model with long context (for example, supporting up to 256k), and reply information is obtained. In this example, 30 text blocks can be selected from each of the two recall methods, for a total of 60 text blocks, greatly increasing the number of input text blocks.
[0033] In this embodiment, training is carried out using long text data, the rotation angle base of the rotary position encoding is increased, and the position of the input text sequence length of the large language model is expanded, thereby increasing the input text sequence length of the large language model and improving the answer integrity and accuracy of the large language model in the problems of long text dependence and multi-document comparison dependence. Further, the text splicing is carried out by using the multi-way recall positive and negative sorting method, removing the sorting model, making the RAG system more concise and efficient and the effect more accurate. In addition, by inserting placeholders, the large language model can more easily distinguish the content at different positions, improving the accuracy of the answer of the large language model.
[0034] The embodiments of the present application provide a training method for a large language model. The method in this embodiment can be applied to servers, terminal devices, platforms, devices, etc. with computing and processing capabilities. Among them, the server can be a server cluster or a single server, and can be a server deployed in the cloud or a local server.
[0035] As Figure 3 shown in the flowchart of the training method for a large language model according to an embodiment of the present application, it includes:
[0036] Step S301: Obtain long text training data, where the sequence length of the long text training data is greater than the maximum length of the input text sequence of the pre-trained large language model.
[0037] Among them, a large language model (LLM) refers to a deep learning model trained using a large amount of text data, which can generate natural language text or understand the meaning of language text. The specific structure of the large language model can be a structure based on the Transformer model.
[0038] Currently, the sequence length (the number of input tokens) of the input text of the open-source pre-trained large language model is generally 8k - 32k. Collect training corpora of long texts, such as data from network datasets or book document data, which includes data in multiple fields, and then clean them to a preset length, such as 128k (k = 1024).
[0039] Step S302: Increase the rotation angle base of the rotary position encoding of the pre-trained large language model to obtain a modified pre-trained large language model.
[0040] Modifying the rotation angle base of the rotary position encoding of the pre-trained large language model, that is, the base value of the rotary position encoding, which is a hyperparameter of the large language model. The rotary position encoding generates position encoding vectors through mathematical transformations. Increasing the rotation angle base of the rotary position encoding can change the way of mathematical transformation, thereby affecting the generation of position encoding vectors, enabling the trained large language model to still effectively capture position information when processing longer sequences, and thus allowing the large language model to process longer input text sequences. For example, change the base value of the rotary position encoding from 10000 to 1000000.
[0041] Step S303: Use the long text training data to train the modified pre-trained large language model to obtain a trained large language model.
[0042] Using long text training data, train the modified pre-trained large language model. Optionally, unsupervised training can be adopted until the training end condition is met, and a trained large language model is obtained. At this time, the maximum length of the input text sequence of the obtained large language model is greater than the maximum length of the input text sequence of the open-source pre-trained large language model.
[0043] The training method of the large language model provided by the embodiments of the present application includes: obtaining long text training data, where the sequence length of the long text training data is greater than the maximum length of the input text sequence of the pre-trained large language model; increasing the rotation angle base of the rotary position encoding of the pre-trained large language model to obtain a modified pre-trained large language model; using the long text training data to train the modified pre-trained large language model to obtain a trained large language model. In this embodiment, by obtaining long text training data and increasing the rotation angle base of the rotary position encoding, the pre-trained large language model is trained. Since the rotary position encoding generates position encoding vectors through mathematical transformations, increasing the rotation angle base of the rotary position encoding can change the way of mathematical transformation, thereby affecting the generation of position encoding vectors, enabling the trained large language model to still effectively capture position information when processing longer sequences, thus allowing the large language model to process longer input text sequences, achieving the amplification of the length of the input text sequence, and improving the answer integrity and accuracy of the large language model in the problems of long text dependence and multi-document comparison dependence.
[0044] In one implementation, the training method of the large language model further includes: obtaining long text training data in a preset domain, and using the long text training data in the preset domain to train the trained large language model to obtain a first large language model corresponding to the preset domain.
[0045] In practical applications, the large language model trained using data from multiple domains is a general large language model. The general large language model learns the general laws and extensive knowledge of language. To improve the accuracy of the large language model in processing information in a preset domain, the large language model is fine-tuned using the long text data in the preset domain.
[0046] Among them, the preset domain can be determined according to the specific domain where the large language model is applied.
[0047] In this embodiment, fine-tuning the large language model using the long text training data in the preset domain can improve the integrity and accuracy of the large language model in processing the long text data in the preset domain.
[0048] In one implementation, the training method of the large language model further includes: in the configuration file of the first large language model corresponding to the preset domain, increasing the maximum position number of the first large language model to obtain a second large language model, and the maximum length of the input text sequence of the second large language model is greater than the maximum length of the input text sequence of the first large language model.
[0049] In practical applications, before using the language model for inference, the sequence length of the input data of the large language model can be further expanded. Specifically, in the configuration file of the large language model corresponding to the preset domain, increase the maximum position number of the large language model (the maximum position number represents the maximum text sequence length that the large language model can process), that is, the value of the parameter original_max_position_embeddings, and increase the value of this parameter to a preset multiple. For example, increase it from 32k to 32k * 8 = 256k.
[0050] In this embodiment, by expanding the position of the input text sequence length of the large language model in the configuration file, the maximum length of the input text sequence of the large language model is further increased, improving the integrity and accuracy of the answers of the large language model in the problems of long text dependence and multi-document comparison dependence.
[0051] The embodiments of the present application provide a training method for a large language model. The method in this embodiment can be applied to servers, terminal devices, platforms, devices, etc. with computing and processing capabilities. Among them, the server can be a server cluster or a single server, and can be a server deployed in the cloud or a local server.
[0052] Such as Figure 4 shown in the flowchart of the training method of the large language model according to an embodiment of the present application, including:
[0053] Step S401, obtain long text training data, and the sequence length of the long text training data is greater than the maximum length of the input text sequence of the pre-trained large language model.
[0054] Currently, the sequence length (the number of input tokens) of the input text of the open-source pre-trained large language model is generally 8k - 32k. Collect the training corpus of long text, such as the data of network datasets or book document data, which includes data in multiple fields, and then clean it to the preset length, such as 128k (k = 1024).
[0055] Step S402, increase the rotation angle base number of the rotation position encoding of the pre-trained large language model to obtain a modified pre-trained large language model.
[0056] Modify the rotation angle base number of the rotary position encoding of the pre-trained large language model, that is, the base value of the rotary position encoding, which is a hyperparameter of the large language model. Increasing the base value of the rotary position encoding can increase the length of the input text sequence of the large language model.
[0057] Step S403: Use the long text training data to train the modified pre-trained large language model to obtain a trained large language model.
[0058] Use the long text training data to train the modified pre-trained large language model. Optionally, the training can be performed in an unsupervised manner until the training end condition is met to obtain a trained large language model.
[0059] Step S404: Obtain the long text training data of the preset domain, and use the long text training data of the preset domain to train the trained large language model to obtain the first large language model corresponding to the preset domain.
[0060] The large language model trained using data from multiple domains is a general large language model. The general large language model learns the general laws and extensive knowledge of language. To improve the accuracy of the large language model in processing information in the preset domain, the large language model is fine-tuned using the long text data of the preset domain.
[0061] Step S405: In the configuration file of the first large language model corresponding to the preset domain, increase the maximum number of positions of the first large language model to obtain the second large language model.
[0062] Among them, the maximum length of the input text sequence of the second large language model is greater than the maximum length of the input text sequence of the first large language model.
[0063] In this embodiment, by obtaining the long text training data and increasing the rotation angle base number of the rotary position encoding, the pre-trained large language model is trained to achieve the amplification of the length of the input text sequence. In this way, the trained large language model can process long text sequences, improving the answer integrity and accuracy of the large language model in the problems of long text dependence and multi-document comparison dependence. Moreover, using the long text training data of the preset domain to fine-tune the large language model can improve the integrity and accuracy of the large language model in processing the long text data of the preset domain.
[0064] The embodiments of the present application provide a question and answer method based on a large language model. The method in this embodiment can be applied to servers, terminal devices, platforms, devices, etc. with computing and processing capabilities. Among them, the server can be a server cluster or a single server, and can be a server deployed in the cloud or a local server.
[0065] Such asFigure 5 The following is a flowchart of a question-answering method based on a large language model according to an embodiment of the present application, including:
[0066] Step S501: Obtain question information and a task instruction, and query multiple text blocks related to the question information.
[0067] Among them, the question information can be a question input by the user or question information called from a preset storage area. The large language model needs to perform reasoning on the question information to obtain a reply information corresponding to the question information.
[0068] The task instruction (Instruction) is a guidance or command required to direct the large language model to perform a specific task, indicating the specific task that the large language model user wants to complete or the type of output required. For example, you are an xxx expert. Please answer the user's question xxx according to the retrieved xxx and in accordance with the xxx requirements.
[0069] Multiple text blocks related to the question information can be queried in different databases through various recall methods. Each text block corresponds to a relevance score, which represents the degree of relevance to the question information. The higher the score, the higher the degree of relevance.
[0070] Among them, for relevance calculation, methods such as term frequency-inverse document frequency (TF-IDF) and text matching algorithm bm25 can be used. The text is tokenized and then the relevance score is calculated.
[0071] Step S502: Concatenate the question information, the task instruction, and the multiple text blocks to obtain a concatenated text.
[0072] Among them, the length of the concatenated text is greater than the maximum length of the input text sequence of the open-source pre-trained large language model.
[0073] Step S503: Input the concatenated text into the trained large language model to obtain the reply information output by the trained large language model.
[0074] Among them, the trained large language model is trained by using the training method in the above embodiment. Since the large language model in this embodiment is trained by obtaining long text training data, increasing the rotation angle base of the rotary position encoding, and performing position expansion on the input text sequence length of the large language model, the trained large language model can process long text data.
[0075] The question-answering method based on a large language model provided by an embodiment of the present application obtains question information and a task instruction, and queries a plurality of text blocks related to the question information; splices the question information, the task instruction, and the plurality of text blocks to obtain a spliced text; inputs the spliced text into a trained large language model to obtain reply information output by the trained large language model; wherein, the trained large language model is trained by using the training method in the embodiment of the present application. In this embodiment, the trained large language model is trained by obtaining long text training data and increasing the rotation angle base of the rotary position encoding, which can realize the amplification of the length of the input text sequence and improve the answer integrity and accuracy of the large language model in the problems of long text dependence and multi-document comparison dependence.
[0076] In one implementation, splicing the question information, the task instruction, and the plurality of text blocks to obtain a spliced text includes: sorting the plurality of text blocks according to the relevance score with the question information, and adding the sorted plurality of text blocks between the task instruction and the question information to obtain a spliced text.
[0077] Wherein, each text block corresponds to a relevance score, which represents the degree of relevance to the question information, sorts the text blocks in descending order according to the relevance score, and splices the sorted text blocks with the task instruction and the question information, and splices them in the order of task execution, sorted text blocks, and question information to obtain a spliced text, and inputs it into the large language model to obtain reply information.
[0078] In this embodiment, when constructing the spliced text, sorting the plurality of text blocks can improve the accuracy of the output result of the large language model. Moreover, sorting according to the relevance score does not require a sorting model compared with the related technology, which can simplify the system structure.
[0079] In one implementation, sorting the plurality of text blocks according to the relevance score with the question information and adding the sorted plurality of text blocks between the task instruction and the question information to obtain a spliced text includes: the plurality of text blocks include a plurality of first text blocks retrieved by the text keyword retrieval method and a plurality of second text blocks retrieved by the vector retrieval method, sorting the plurality of first text blocks in descending order according to the relevance score with the question information, and sorting the plurality of second text blocks in ascending order according to the relevance score with the question information; adding the sorted plurality of first text blocks and the sorted plurality of second text blocks between the task instruction and the question information to obtain a spliced text.
[0080] Retrieve a plurality of text blocks by the vector retrieval method (querying in a vector library), and retrieve a plurality of text blocks by the text keyword (querying by using a search engine library, for example, Elasticsearch library) retrieval method.
[0081] Since large language models are more sensitive to text closer to the task instruction and the part closer to the question information, and since the text chunks obtained by vector recall are more accurate than those obtained by text keyword recall, therefore, the text chunks obtained by text keyword recall are sorted from high to low according to the relevance score with the question information, and the text chunks obtained by vector recall are sorted from low to high according to the relevance score with the question information. Then, the text chunks obtained by text keyword recall are concatenated after the task instruction, and the text chunks obtained by vector recall are concatenated before the question information to obtain the concatenated text. After the concatenated text is input into the large language model, the accuracy of the output result can be improved.
[0082] In one implementation, after concatenating the question information, task instruction, and multiple text chunks to obtain the concatenated text, the question-answering method based on the large language model further includes: adding a first placeholder after the task instruction in the concatenated text, and adding a second placeholder before the question information in the concatenated text to obtain the processed concatenated text. The first placeholder is used to indicate the position corresponding to the task instruction, and the second placeholder is used to indicate the position corresponding to the question information.
[0083] To make it easier for the large language model to distinguish the content at different positions (task instruction, text chunk, question information), a first placeholder Pre-Tokens and a second placeholder Post-Tokens are respectively inserted. The number of placeholders for different large language models is different. For example, Pre-Tokens can be 400, and Post-Tokens can be 100. Finally, the concatenated text is input to the large language model with long context (for example, supporting up to 256k), and the reply information is obtained.
[0084] In this embodiment, by inserting placeholders, it is easier for the large language model to distinguish the content at different positions, and the accuracy of the large language model's answer is improved.
[0085] The embodiment of the present application provides a question-answering method based on a large language model. The method in this embodiment can be applied to servers, terminal devices, platforms, devices, etc. with computing and processing capabilities. Among them, the server can be a server cluster or a single server, and can be a server deployed in the cloud or a local server.
[0086] As Figure 6 shown is the flowchart of the question-answering method based on the large language model according to an embodiment of the present application, including:
[0087] Step S601, obtain the question information and the task instruction.
[0088] Step S602, query multiple first text chunks related to the question information according to the text keyword recall method.
[0089] Step S603, query multiple second text chunks related to the question information according to the vector recall method.
[0090] Step S604, sort the multiple first text chunks in descending order according to the relevance score with the question information, and sort the multiple second text chunks in ascending order according to the relevance score with the question information.
[0091] Step S605, add the sorted multiple first text chunks and the sorted multiple second text chunks between the task instruction and the question information to obtain a concatenated text.
[0092] Wherein, the sequence length of the concatenated text is greater than the maximum length of the input text sequence of the pre-trained large language model.
[0093] Step S606, add a first placeholder after the task instruction in the concatenated text, and add a second placeholder before the question information in the concatenated text to obtain a processed concatenated text.
[0094] Wherein, the first placeholder is used to indicate the position corresponding to the task instruction, and the second placeholder is used to indicate the position corresponding to the question information.
[0095] Step S607, input the processed concatenated text into the trained large language model to obtain the reply information output by the trained large language model.
[0096] Wherein, the trained large language model is trained by using the training method in the above embodiment.
[0097] In this embodiment, the trained large language model is trained by obtaining long text training data and increasing the rotation angle base of the rotary position encoding, and the position of the input text sequence length of the trained large language model is expanded, so as to realize the amplification of the length of the input text sequence and improve the answer integrity and accuracy of the large language model in the problems of long text dependence and multi-document comparison dependence.
[0098] Corresponding to the application scenario and method of the method provided in the embodiment of the present application, the embodiment of the present application also provides a training device for a large language model. As Figure 7 Shown is the structural block diagram of the training device for a large language model according to an embodiment of the present application. The device includes:
[0099] A data acquisition module 701, configured to acquire long text training data, and the sequence length of the long text training data is greater than the maximum length of the input text sequence of the pre-trained large language model.
[0100] A parameter modification module 702, configured to increase the rotation angle base number of the rotary position encoding of the pre-trained large language model to obtain a modified pre-trained large language model.
[0101] A model training module 703, configured to train the modified pre-trained large language model by using long text training data to obtain a trained large language model.
[0102] The training device for the large language model provided by the embodiment of the present application obtains long text training data, and the sequence length of the long text training data is greater than the maximum length of the input text sequence of the pre-trained large language model; increases the rotation angle base number of the rotary position encoding of the pre-trained large language model to obtain a modified pre-trained large language model; and trains the modified pre-trained large language model by using the long text training data to obtain a trained large language model. In this embodiment, by obtaining long text training data and increasing the rotation angle base number of the rotary position encoding, the pre-trained large language model is trained. Since the rotary position encoding generates position encoding vectors through mathematical transformations, increasing the rotation angle base number of the rotary position encoding can change the way of mathematical transformation, thereby affecting the generation of position encoding vectors, so that the trained large language model can still effectively capture position information when processing longer sequences, thus allowing the large language model to process longer input text sequences, achieving the amplification of the length of the input text sequence, and improving the answer integrity and accuracy of the large language model in the problems of long text dependence and multi-document comparison dependence.
[0103] In one implementation, the training device for the large language model is further configured to: obtain long text training data in a preset domain, and train the trained large language model by using the long text training data in the preset domain to obtain a first large language model corresponding to the preset domain.
[0104] In one implementation, the training device for the large language model is further configured to: increase the maximum number of positions of the first large language model in the configuration file of the first large language model corresponding to the preset domain to obtain a second large language model, and the maximum length of the input text sequence of the second large language model is greater than the maximum length of the input text sequence of the first large language model.
[0105] The functions of the modules in the embodiments of the present application can refer to the corresponding descriptions in the above methods, and have corresponding beneficial effects, which will not be elaborated here.
[0106] Corresponding to the application scenario and method of the method provided by the embodiment of the present application, the embodiment of the present application further provides a question and answer device based on a large language model. As Figure 8 shown is the structural block diagram of the question and answer device based on the large language model according to an embodiment of the present application. The device includes:
[0107] An information acquisition module 801, configured to acquire problem information and task instructions, and query multiple text blocks related to the problem information.
[0108] A text splicing module 802, configured to splice the problem information, task instructions, and multiple text blocks to obtain a spliced text; the sequence length of the spliced text is greater than the maximum length of the input text sequence of the pre-trained large language model.
[0109] An information processing module 803, configured to input the spliced text into the trained large language model to obtain the reply information output by the trained large language model; wherein, the trained large language model is trained by using the training method in the embodiments of the present application.
[0110] The question-answering device based on a large language model provided by the embodiments of the present application acquires problem information and task instructions, and queries multiple text blocks related to the problem information; splices the problem information, task instructions, and multiple text blocks to obtain a spliced text; the sequence length of the spliced text is greater than the maximum length of the input text sequence of the pre-trained large language model; inputs the spliced text into the trained large language model to obtain the reply information output by the trained large language model; wherein, the trained large language model is trained by using the training method in the embodiments of the present application. In this embodiment, the trained large language model is trained by acquiring long text training data and increasing the rotation angle base number of the rotary position encoding, which can realize the amplification of the length of the input text sequence and improve the answer integrity and accuracy of the large language model in the problems of long text dependence and multi-document comparison dependence.
[0111] In one implementation, the text splicing module 802 is configured to: sort the multiple text blocks according to the relevance scores with the problem information, and add the sorted multiple text blocks between the task instructions and the problem information to obtain a spliced text.
[0112] In one implementation, the text splicing module 802 is specifically configured to: the multiple text blocks include multiple first text blocks queried by the text keyword recall method and multiple second text blocks queried by the vector recall method, sort the multiple first text blocks from high to low according to the relevance scores with the problem information, and sort the multiple second text blocks from low to high according to the relevance scores with the problem information; add the sorted multiple first text blocks and the sorted multiple second text blocks between the task instructions and the problem information to obtain a spliced text.
[0113] In one implementation, the question-and-answer device based on the large language model is further configured to: after splicing the question information, the task instruction, and multiple text blocks to obtain a spliced text, add a first placeholder after the task instruction in the spliced text, and add a second placeholder before the question information in the spliced text to obtain a processed spliced text, where the first placeholder is used to indicate the position corresponding to the task instruction, and the second placeholder is used to indicate the position corresponding to the question information.
[0114] For the functions of the modules in the embodiments of the present application, reference may be made to the corresponding descriptions in the above methods, and the corresponding beneficial effects are achieved, which will not be elaborated here.
[0115] Figure 9 FIG. is a block diagram of an electronic device for implementing the embodiments of the present application. As Figure 9 shown, the electronic device includes: a memory 910 and a processor 920. The memory 910 stores a computer program that can run on the processor 920. When the processor 920 executes the computer program, the method in the above embodiments is implemented. The number of the memory 910 and the processor 920 may be one or more.
[0116] The electronic device further includes:
[0117] A communication interface 930, configured to communicate with external devices and perform data interaction and transmission.
[0118] If the memory 910, the processor 920, and the communication interface 930 are implemented independently, the memory 910, the processor 920, and the communication interface 930 may be interconnected through a bus and communicate with each other. The bus may be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus, etc. The bus may be divided into an address bus, a data bus, a control bus, etc. For the sake of representation, Figure 9 only a thick line is used to represent it in the figure, but it does not mean that there is only one bus or one type of bus.
[0119] Optionally, in a specific implementation, if the memory 910, the processor 920, and the communication interface 930 are integrated on a chip, the memory 910, the processor 920, and the communication interface 930 may communicate with each other through an internal interface.
[0120] An embodiment of the present application provides a computer-readable storage medium storing a computer program, and when the program is executed by a processor, the method provided in the embodiment of the present application is implemented.
[0121] An embodiment of the present application provides a computer program product, which includes a computer program, and when the computer program is executed by a processor, the method provided in the embodiment of the present application is implemented.
[0122] An embodiment of the present application further provides a chip, which includes a processor for calling and running instructions stored in a memory from the memory, so that a communication device installed with the chip executes the method provided in the embodiment of the present application.
[0123] An embodiment of the present application further provides a chip, including: an input interface, an output interface, a processor, and a memory. The input interface, the output interface, the processor, and the memory are connected through an internal connection path. The processor is used to execute the code in the memory, and when the code is executed, the processor is used to execute the method provided in the embodiment of the application.
[0124] It should be understood that the above-mentioned processor may be a central processing unit (CPU), or may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor, etc. It is worth noting that the processor may be a processor supporting the advanced reduced instruction set machine (ARM) architecture.
[0125] Further, optionally, the above-mentioned memory may include a read-only memory and a random access memory. The memory may be a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memories. Among them, the non-volatile memory may include a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory may include a random access memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of RAM are available. For example, static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM), and direct rambus random access memory (DR RAM).
[0126] In the above embodiments, it may be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it may be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions according to the present application are generated in whole or in part. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions may be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium.
[0127] In the description of this specification, the descriptions referring to terms such as "one embodiment", "some embodiments", "example", "specific example", or "some examples", etc., mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of this application. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.
[0128] In addition, the terms "first" and "second" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include at least one of such features. In the description of this application, "a plurality of" means two or more unless otherwise specifically defined.
[0129] Any process or method described in the flowchart or otherwise described herein can be understood to represent a module, segment, or portion of code including one or more executable instructions for implementing a specific logical function or process. And the scope of the preferred embodiments of this application includes additional implementations, where the functions can be executed in a substantially simultaneous manner or in a reverse order according to the involved functions, rather than in the order shown or discussed.
[0130] The logic and / or steps described in the flowchart or otherwise described herein, for example, can be considered as a sequenced list of executable instructions for implementing a logical function, and can be specifically implemented in any computer-readable medium for use by an instruction execution system, apparatus, or device (such as a computer-based system, a system including a processor, or other systems that can fetch and execute instructions from the instruction execution system, apparatus, or device), or in conjunction with these instruction execution systems, apparatus, or devices.
[0131] It should be understood that each part of this application can be implemented by hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. All or part of the steps of the method in the above embodiments can be completed by a program instructing relevant hardware, and this program can be stored in a computer-readable storage medium. When this program is executed, it includes one or a combination of the steps of the method embodiment.
[0132] In addition, each functional unit in various embodiments of the present application may be integrated into a processing module, may exist separately as individual physical units, or two or more units may be integrated into one module. The above-mentioned integrated module may be implemented in the form of hardware or in the form of a software functional module. When the above-mentioned integrated module is implemented in the form of a software functional module and sold or used as an independent product, it may also be stored in a computer-readable storage medium. The storage medium may be a read-only memory, a magnetic disk, an optical disc, or the like.
[0133] As described above, only the exemplary embodiments of the present application are provided, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope recorded in the present application can easily think of various changes or substitutions, and these should all be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
Claims
1. A question-answering method based on a large language model, characterized in that: include: Obtaining problem information and task instructions, and querying multiple text blocks related to the problem information; The question information, the task instruction and the multiple text blocks are concatenated to obtain a concatenated text; the sequence length of the concatenated text is greater than the maximum length of the input text sequence of the pre-trained large language model; Inputting the concatenated text into the trained large language model to obtain reply information output by the trained large language model; The step of splicing the problem information, the task instruction and the multiple text blocks to obtain a spliced text includes: sorting the multiple text blocks according to the relevance scores with the question information, and adding the sorted multiple text blocks between the task instruction and the question information to obtain a concatenated text; The step of sorting the multiple text blocks according to the relevance scores with the question information, and adding the sorted multiple text blocks between the task instruction and the question information to obtain a concatenated text includes: The multiple text blocks include multiple first text blocks obtained by querying in a text keyword recall manner and multiple second text blocks obtained by querying in a vector recall manner, and the multiple first text blocks are sorted from high to low according to the relevance scores with the question information, and the multiple second text blocks are sorted from low to high according to the relevance scores with the question information; Adding the sorted multiple first text blocks and the sorted multiple second text blocks between the task instruction and the question information to obtain a concatenated text; The trained large language model is obtained by training in the following way: Acquire long text training data, where the sequence length of the long text training data is greater than the maximum length of an input text sequence of a pre-trained large language model; Increasing the rotation angle base of the rotation position encoding of the pre-trained large language model to obtain a modified pre-trained large language model; The modified pre-trained large language model is trained using the long text training data to obtain a trained large language model.
2. The method according to claim 1, characterized in that After the question information, the task instruction and the plurality of text blocks are spliced together to obtain a spliced text, the method further includes: A first placeholder is added after the task instruction in the spliced text, and a second placeholder is added before the question information in the spliced text to obtain a processed spliced text, wherein the first placeholder is used to indicate a position corresponding to the task instruction, and the second placeholder is used to indicate a position corresponding to the question information.
3. An electronic device, characterized in that: The electronic device comprises a memory, a processor and a computer program stored in the memory, and the processor implements the method according to any one of claims 1 to 2 when executing the computer program.
4. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1 to 2 is implemented.
5. A computer program product, characterized in that The computer program product comprises a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1 to 2 is implemented.
Citation Information
Patent Citations
Retrieval question and answer method, system and equipment and medium
CN117708309A