Large model LLaMA3 fine-tuning knowledge question and answer dialogue system based on local knowledge base
By adopting the LLaMA3 fine-tuning technology based on the local knowledge base in the Q&A system, the accuracy and cost problems of traditional Q&A systems when understanding natural language and building knowledge graphs are solved, and more efficient and accurate Q&A capabilities are achieved.
Patent Information
- Application Number
- CN202510105477.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-23
- Publication Date
- 2025-05-30
AI Technical Summary
Traditional question-and-answer systems have limitations on accuracy and scope of application when understanding natural language, and the cost of building and maintaining large-scale knowledge graphs is high, making it difficult for the system to reason and generate accurate answers, and face the problems of performance bottlenecks and poor scalability.
The large model LLaMA3 fine-tuning knowledge question and answer dialogue system based on the local knowledge base is adopted to achieve accurate question and answer text generation by obtaining the parameters of the LLaMA3 large model, building a low-rank decomposition technology model, training the model, building a local knowledge base, fine-tuning the model, evaluating the model and deploying the model.
It reduces the complex high-cost pre-training of large models, quickly obtains better model results, improves the accuracy, fluency and long-sequence processing capabilities of the Q&A system, and solves the development cost and performance bottlenecks of traditional systems.
Smart Images

Figure CN120067249A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of knowledge - based question - answering dialogue, and particularly to a fine - tuned knowledge - based question - answering dialogue system of the large - model LLaMA3 based on a local knowledge base. Background Art
[0002] As an artificial intelligence system capable of answering user questions, the background technical solutions of the question - answering system involve multiple key technical fields. The construction of the question - answering system mainly relies on natural language processing (NLP), machine learning, knowledge graphs, database technology, and web development technology. The comprehensive application of these technologies enables the question - answering system to understand the questions raised by users and retrieve relevant information from the knowledge base to generate accurate answers. Natural language processing technology is used to process user questions, including steps such as word segmentation, part - of - speech tagging, syntactic analysis, and semantic understanding. This helps to convert the natural - language questions of users into a form that can be understood by a computer and retrieve relevant answers from the knowledge base. Machine learning technology plays an important role in the question - answering system and is used for tasks such as question classification, information retrieval, and matching. By training the model, the system can more accurately understand user intentions and improve the accuracy and efficiency of answers. Knowledge graph technology is used to construct and manage the knowledge base, represent knowledge in the form of a graph, and provide query and reasoning functions. This enables the question - answering system to retrieve and integrate information more efficiently and generate more accurate answers. The database is used to store and manage the data of the question - answering system, including user questions, system answers, knowledge - base content, etc. An efficient data - storage and retrieval mechanism is the basis for the stable operation of the question - answering system. Web development technology includes technologies such as HTML, CSS, and JavaScript, which are used to build the user interface and realize the interaction between the front - end and back - end of the system. A friendly user interface and a smooth user experience are the keys to attracting users for the question - answering system.
[0003] In recent years, neural - network question - answering systems have certain advantages compared to probabilistic - statistical models. The Word Embedding technology is used to convert text into a representation in a high - dimensional vector space. Models such as convolutional neural networks (CNNs), recurrent neural networks (RNNs), or their variants (such as LSTMs, GRUs) are used to extract the semantic features of questions. When it comes to multi - turn conversations or document reading comprehension, it is also necessary to extract context features to understand the complete context of the question. Technologies such as the Attention Mechanism are used to match the question with candidate answers, calculate the similarity or correlation between them, and for questions that directly extract answers from the text, perform grammar checking, fluency optimization, etc. on the generated answers to improve the user experience.
[0004] The main drawbacks are as follows:
[0005] The methods of text generation and dialogue in traditional neural networks are limited by defects such as sequence length and insufficient parallel computing. Question-and-answer systems are difficult to accurately understand natural language, especially when the user's questions are complex or ambiguous. This limits the accuracy of the system's answers and the scope of application. Building a moderately large knowledge graph requires a large amount of human and material resources, including steps such as data sample extraction, expert annotation, and machine learning training. The high cost makes it difficult to popularize question-and-answer systems. Question-and-answer systems may have difficulty reasoning and generating accurate answers from the knowledge graph. This requires the system to have more advanced reasoning capabilities and a richer knowledge base support. With the increase in the number of users and the growth of data volume, question-and-answer systems may face performance bottleneck problems. This requires the system to have efficient data processing capabilities and good scalability to meet the growing user needs.
[0006] Relatively speaking, in the question-and-answer systems of traditional models, the communication and response between users and the system are slow, the question-and-answer systems are relatively rigid and inflexible, and the expansion flexibility is poor. In recent years, popular large models are based on huge datasets and high computing power consumption, and small companies and relevant service departments in the field cannot directly deploy them due to high costs. At the same time, for traditional field medical diagnosis and consultation question-and-answer systems and the like, the model training and deployment processes are cumbersome and the systems are complex. Based on the fact that the user objects are very serious and the credibility requirements of the model results are extremely high, it is often necessary to repeatedly modify and optimize during model training and deployment.
[0007] Therefore, there is an urgent need in this field for a technical solution that can solve the above defects of traditional question-and-answer systems.
[0008] The information disclosed in this background art section is only intended to increase the understanding of the overall background of the present invention and should not be regarded as an admission or any form of implication that this information constitutes the prior art already known to those of ordinary skill in the art. Summary of the Invention
[0009] The object of the present invention is to provide a technical solution that can solve the defects such as high development cost of traditional question-and-answer systems.
[0010] To achieve the above object, the present invention provides the following solution:
[0011] A large model LLaMA3 fine-tuning knowledge question-and-answer dialogue system based on a local knowledge base, comprising:
[0012] An LLaMA3 large model parameter acquisition module for acquiring LLaMA3 large model parameters;
[0013] A low-rank decomposition technology construction model establishment module for establishing a low-rank decomposition technology construction model;
[0014] A model training module for training the low-rank decomposition technology construction model;
[0015] Local knowledge base construction module, used to construct a local knowledge base;
[0016] Training model fine-tuning module, used to fine-tune the training model;
[0017] Model evaluation module, used to evaluate the model;
[0018] Model deployment module, used to deploy the model.
[0019] Optionally, the fine-tuned training model includes:
[0020] Construct a dataset;
[0021] Construct a base model and pre-train the LLaMA base model;
[0022] Supervised fine-tuning (SFT) for specific tasks;
[0023] Reinforcement learning from human feedback (RLHF) to improve the model's question-answering quality;
[0024] Data augmentation and optimization of the local knowledge base;
[0025] Model evaluation, deployment, and continuous optimization.
[0026] Optionally, the construction of the dataset is specifically: using a high-precision model or a large model to construct the dataset.
[0027] Optionally, the construction of the base model and pre-training of the LLaMA base model includes:
[0028] Input feature root mean square normalization (RMS Norm), grouped query attention mechanism, and residual connection; then regularize the sample data to ensure that the data control generates smaller errors; then perform another residual connection and normal distribution normalization to complete the calculation of the entire Decoding part.
[0029] Optionally, the specific tasks of the supervised fine-tuning (SFT) are specifically:
[0030] Supervised fine-tuning is to pre-train a source neural network model on a source dataset, and then use the downstream scenario data to create a new target neural network model. The target model replicates all the model designs and most of the parameters of the source model except for the output layer. These model parameters contain the knowledge learned on the source dataset, and this knowledge is also applicable to the target dataset. During fine-tuning, an output layer with an output size equal to the number of classes in the target dataset is added to the target model, and the model parameters of this layer are randomly initialized. When training the target model on the target dataset, it will be trained from scratch to the output layer, and the parameters of the remaining layers are fine-tuned based on the parameters of the source model.
[0031] Optionally, the improvement of the model's question-answering quality by Reinforcement Learning from Human Feedback (RLHF) is specifically as follows:
[0032] Supervised fine-tuning, using normal instruction following or dialogue samples to train the model's basic dialogue and ability to follow prompts; based on human preferences and annotations, training a scoring model RM that can simulate human preferences; using RM to provide feedback and continuously adjusting the model's behavior through the Proximal Policy Optimization (PPO) reinforcement learning framework.
[0033] Optionally, the local knowledge base data enhancement and optimization is specifically as follows:
[0034] Input a prompt retrieval sequence, embed the query sequence into a word vector model, use the query sequence word vectors to perform similarity matching with the vectors in the vector database, retrieve the corresponding documents in the word vector database, find the top k sequences with the highest similarity as the reference sequence vectors for answering, combine the retrieved text sequences with the query sequence to construct a better prompt, and then pass this prompt to the large language model. Given the input sequence X, after word vector embedding, X_embeding = embedding(X, V), where V is the word vector table, the word vectors enter the database for matching, and by calculating the two, we obtain the text X_query = Q(X_embedding, D), and the text output is used as the input series for the large model.
[0035] Optionally, the model evaluation, deployment, and continuous optimization are specifically as follows:
[0036] First, download the llama3 model file. Using the pre-trained model + LORA adapter method, the model can be quickly and efficiently adapted to new tasks or domains without retraining the entire model.
[0037] Characterize the accuracy after model fine-tuning and evaluate the output sequence of the model; after building the local knowledge base and model training, through the parallel computing of the high-performance GPU A10 architecture of the aliyun artificial intelligence platform, the evaluation metrics of the model can be obtained; for model evaluation, select the large-scale multi-task language understanding benchmark for testing. The evaluation method uses the questions in each subject area to calculate the correct rate of the model respectively, and calculate the average correct rate of all subjects as the overall performance of the model, and compare the performance of the model with the correct rate of human experts.
[0038] Optionally, it also includes establishing a large model question-answering application platform.
[0039] Optionally, the establishment of the large model question-answering application platform includes:
[0040] Build a local knowledge base. Input local PDF files through the system to build a machine learning Q&A knowledge base. Upload a large number of machine learning PDF files, and then through the file processing component, semantically understand, segment, and vectorize the file data and store it in the vector database and search engine. Synchronously associate the files with the knowledge base, and store the associated information in the MySQL database to build the knowledge base, and perform training and inference on the large model Q&A interface.
[0041] Compared with the prior art, the present invention has the following beneficial effects:
[0042] The device provided by the present invention can collect the local knowledge base according to the large model schema, construct large model fine-tuning, train, test, and evaluate the model through the local knowledge base, and realize the generation of accurate Q&A texts for the local knowledge base, thereby reducing the complex and high-cost model pre-training of the large model and quickly obtaining better model results.
[0043] The present invention can solve the defects of traditional Q&A systems. Based on the powerful model expression ability of the large model, through the local knowledge base, an accurate model is fine-tuned, and it has great advantages in Q&A accuracy, fluency, and long sequences. BRIEF DESCRIPTION OF THE DRAWINGS
[0044] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0045] Figure 1 It is a flowchart of the transformer provided for the embodiment of the present invention.
[0046] Figure 2 It is a framework diagram of LLaMMa3 provided for the embodiment of the present invention.
[0047] Figure 3 It is an improved block diagram of LLaMMa3 provided for the embodiment of the present invention.
[0048] Figure 4 It is a schematic diagram of the method for constructing a local knowledge base provided for the embodiment of the present invention.
[0049] Figure 5 It is a schematic diagram of the attention mechanism and relative position embedding of LLaMMa3 provided for the embodiment of the present invention.
[0050] Figure 6 It is a schematic diagram of root mean square normalization and residual connection provided for the embodiment of the present invention.
[0051] Figure 7 Schematic diagram of the self-attention module provided by an embodiment of the present invention.
[0052] Figure 8 Schematic diagram of LoRa low-rank decomposition and weight calculation provided by an embodiment of the present invention.
[0053] Figure 9 Block diagram of local knowledge base data enhancement provided by an embodiment of the present invention.
[0054] Figure 10 Block diagram of the evaluation of the LLaMMa3 large model provided by an embodiment of the present invention.
[0055] Figure 11 Schematic diagram of constructing a local knowledge base provided by an embodiment of the present invention.
[0056] Figure 12 Schematic diagram of large model question answering in the local knowledge base provided by an embodiment of the present invention. Detailed implementation manners
[0057] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0058] The purpose of the present invention is to provide a technical solution that can solve the defects such as high development cost of traditional question answering systems.
[0059] To make the above objects, features, and advantages of the present invention more obvious and understandable, the present invention will be further described in detail below in conjunction with the accompanying drawings and specific implementation manners.
[0060] Embodiment 1:
[0061] This embodiment provides a large model LLaMA3 fine-tuning knowledge question answering dialogue system based on a local knowledge base, including:
[0062] LLaMA3 large model parameter acquisition module, used to acquire LLaMA3 large model parameters;
[0063] Low-rank decomposition technology model establishment module, used to establish a low-rank decomposition technology model;
[0064] Model training module, used to train the low-rank decomposition technology model;
[0065] Local knowledge base construction module, used to construct a local knowledge base;
[0066] Training model fine-tuning module, used to fine-tune the training model;
[0067] Model evaluation module, used to evaluate the model;
[0068] Model deployment module, used to deploy the model.
[0069] The LLaMMA3 large model is based on the Decoder of the transformers module as Figure 1 , and is built through module modification into the Llama large model structure as Figure 2 .
[0070] The following gives the extended block diagram of the model framework Figure 2 , as Figure 3 shown.
[0071] The following gives the steps for fine-tuning based on the LLaMA3 large model:
[0072] a) Dataset construction
[0073] b) Build the base model and pre-train the LLaMA base model.
[0074] c) Supervised fine-tuning of specific tasks for SFT.
[0075] d) Reinforcement learning based on human feedback RLHF to improve the model's question-answering quality.
[0076] e) Data enhancement and optimization of the local knowledge base.
[0077] f) Model evaluation, deployment, and continuous optimization.
[0078] Specifically:
[0079] (1) Construct a dataset. Usually, a high-precision model or a large model is used to construct the dataset. In this project, the dataset is generated by manually constructing the dataset and combining it with tools for automatically generating the dataset, and data preprocessing is performed on the data. Through modification, streamlining, and data augmentation, high-quality data is obtained.
[0080] The local knowledge base of this model is built through the following three methods, as Figure 4 .
[0081] The local knowledge base is constructed by reading text pdfs to build json file formats, which are constructed in the form of questions and answers. Local file annotation is used to manually construct relatively accurate question-and-answer form data. Some datasets are generated with the help of relevant model construction. Some question-and-answer data files are generated using tools for automatically generating datasets.
[0082] The data of the fine-tuned local knowledge base must be properly processed to enhance data quality and clean redundant data, with a total of 30,000 pieces of data. The format of the processed data is as follows:
[0083]
[0084]
[0085] (2) Build a base model and pre-train the LLaMA base model.
[0086] The root mean square normalization (RMSNorm) of the input features can simply and efficiently process data. It has obvious advantages especially for long sequence scenarios. The grouped query attention mechanism is used with residual connections as Figure 5 shown.
[0087] Through the word vector model word2vector, each feature of the corpus X is embedded into a word vector by looking up the table
[0088] X = RMSNorm(embedding - lookup(X)),
[0089] For each token, it is necessary to obtain the context semantic information and also embed the weights of other tokens that each token pays attention to. Therefore, it is necessary to calculate the attention scores between each token, especially the attention component of the current token relative to other features.
[0090] Calculate the attention vector coefficients of each token. First, calculate the query matrix and the key-value matrix.
[0091] Q = (RoPE * f(Xm, m), K = RoPE * g(Xn, n), V = h(X) = XW V ,
[0092] The three matrices are related to the attention information between tokens. At the same time, the rotary position matrix RoPE is added, and relative position encoding is adopted.
[0093] X attention = Self - Attention(Q) T *(K)) / sqr(d) * V
[0094] Obtain the contextual semantic information and position information of all tokens associated with each token. Especially in the bidirectional auto-encoding mode, a great deal of information about the context, semantics, and absolute positions between tokens is learned. Through the self-attention model, it can be learned which tokens a token will pay attention to, and features such as whether the attention of this token relative to other tokens is uniform, dispersed, or concentrated. Due to the attenuation of the depth of the network and the fully connected information, a residual connection is added to maintain the network information transmission and the stability of the network architecture.
[0095] X attention = X + X attention
[0096] Then regularize the sample data to ensure that the data control produces smaller errors, as Figure 6 shown.
[0097] X attention = RMSNorm(X attention )
[0098] In the next layer of the network structure, perform forward propagation, perform two-layer linear mapping, and calculate using the activation function as Figure 5 shown.
[0099] X hidden = FFN(SwiGLU(Linear(X attention )))
[0100] Perform a residual connection and normal distribution normalization again
[0101] X hidden = X attention + X hidden
[0102] X hidden = RMSNorm(X hidden )
[0103] Complete the calculation of the entire Decoding part through this architecture as Figure 7 .
[0104] (3) Specific tasks of supervised fine-tuning SFT
[0105] Supervised fine-tuning involves pre-training a source neural network model on a source dataset and then creating a new target neural network model using downstream scenario data. The target model replicates all the model designs and most of the parameters of the source model except for the output layer. These model parameters contain the knowledge learned on the source dataset, and this knowledge is also applicable to the target dataset. During fine-tuning, an output layer with an output size equal to the number of classes in the target dataset is added to the target model, and the model parameters of this layer are randomly initialized. When training the target model on the target dataset, the output layer is trained from scratch, and the parameters of the remaining layers are fine-tuned based on the parameters of the source model.
[0106] During the fine-tuning of large language models, the LoRA low-rank decomposition technique freezes the pre-trained model weights and decomposes the rank of the trainable weight matrix into two low-rank matrices W 0 +ΔW = W 0 +BA, which is added to each layer of the Transformer structure. During LoRA fine-tuning, A is initialized with random Gaussian values and B is initialized with zeros. Therefore, ΔW = BA is zero at the start of training, as Figure 8 shown. Then, the fine-tuned weight parameters are obtained through inference.
[0107] The number of model fine-tuning parameters is greatly reduced. When deployed to a production environment, only W = W 0 +BA needs to be calculated and stored, and inference is performed as usual. Compared with other methods, there is no additional latency because no additional layers need to be added. Fine-tuning the low-rank matrices A and B greatly reduces the total number of model parameters, usually from the billion level to the million level.
[0108] (4) Improving the model's question-answering quality based on human feedback reinforcement learning RLHF
[0109] For supervised fine-tuning, normal instruction following or dialogue samples are used to train the model's basic dialogue and ability to follow prompts; based on human preferences and annotations, a scoring model RM that can simulate human preferences is trained; with the feedback provided by RM, the behavior of the model is continuously adjusted through the PPO reinforcement learning framework.
[0110] (5) Local knowledge base data augmentation and optimization
[0111] Input the prompt retrieval sequence such as Figure 9As shown, the query sequence is embedded into the word vector model, and the similarity between the query sequence word vectors and the vectors in the vector database is matched. The corresponding documents are retrieved from the word vector database, and the top k sequences with the highest similarity are found as the reference sequence vectors for the answer. Therefore, the retrieved text sequence is combined with the query sequence to construct a better prompt, and then this prompt is passed to the large language model. Given the input sequence X, after word vector embedding, X_embeding = embedding(X, V), where V is the word vector table. The word vectors enter the database for matching, and the text X_query = Q(X_embedding, D) is obtained by calculating the similarity between the two. The text output is used as the input series for the large model.
[0112] (6) Model evaluation, deployment, and continuous optimization
[0113] LLaMA3 is a tool for fine-tuning large language models (LLMs). It aims to simplify the fine-tuning process of large language models, enabling users to quickly train and optimize the model to improve its performance on specific tasks.
[0114] First, download the llama3 model file. Using the pre-trained model + LORA adapter method, the model can be quickly and efficiently adapted to new tasks or domains without retraining the entire model. For example, adapting an English model to Chinese conversations is a common and practical fine-tuning method.
[0115] To train the model, first modify the model configuration file llama3——lorao_sft.yaml.
[0116] model
[0117] model_name_or_path: / mnt / workspace / models / Meta-Llama-3-8B-Instruct
[0118] method
[0119] stage:sft do_train:true
[0120] finetuning_type:lora
[0121] lora_target:q_proj,v_proj
[0122] dataset
[0123] dataset:alpaca_gpt4_zh
[0124] template: llama3
[0125] cutoff_len: 1024
[0126] max_samples: 1000o
[0127] verwrite_cache: true
[0128] preprocessing_num_workers: 16
[0129] output
[0130] output_dir: / mnt / workspace / models / llama3 - lora - zh l
[0131] ogging_steps: 100s
[0132] ave_steps: 500
[0133] plot_loss: true
[0134] overwrite_output_dir: true
[0135] train
[0136] per_device_train_batch_size: 1
[0137] gradient_accumulation_steps: 8
[0138] learning_rate: 0.0001
[0139] num_train_epochs: 1.0
[0140] lr_scheduler_type: cosine
[0141] warmup_steps: 0.1
[0142] fp16: true
[0143] eval val_size: 0.1
[0144] per_device_eval_batch_size: 1
[0145] evaluation_strategy: steps
[0146] eval_steps: 500
[0147] Clone the dataset and modify dataset_info.json
[0148] git clone https: / / www.modelscope.cn / datasets / llamafactory / alpaca_gpt4_zh.git
[0149] Modify the dataset_info.json file
[0150] Inference, modify the configuration file llama3_lora_sft.yaml,
[0151] model_name_or_path: / mnt / workspace / models / Meta-Llama-3-8B-Instruct adapter_name_or_path: / mnt / workspace / models / llama3-lora-zh template: llama3 finetuning_type: lora
[0152] Execute the self-supervised fine-tuning file and display the inference interface. Just enter a question at the user position, and the large model will quickly answer the result.
[0153] The evaluation of the model is divided into two parts. The first part is the accuracy characterization after model fine-tuning and the sequence evaluation of the model output results.
[0154] After building the local knowledge base, through model training and parallel computing on the high-performance GPU A10 architecture of the aliyun artificial intelligence platform, the evaluation metrics of the model can be obtained. For model evaluation, a large-scale multi-task language understanding benchmark can be selected for testing, mainly using MMLU (Massive Multitask Language Understanding) to evaluate the model. The evaluation method is to use the questions in each subject area to calculate the correct rate of the model respectively. Calculate the average correct rate of all subjects as the overall performance of the model. Compare the performance of the model with the correct rate of human experts. The configuration file llama3_lora_eval.yaml must be modified during evaluation.
[0155] model
[0156] model_name_or_path: / mnt / workspace / models / Meta-Llama-3-8B-Instruct adapter_name_or_path: / mnt / workspace / models / llama3-lora-zh
[0157] method
[0158] finetuning_type: lora
[0159] dataset task: mmlu
[0160] split: test
[0161] template: fewshot
[0162] lang: en
[0163] n_shot: 5
[0164] output
[0165] save_dir: saves / llama3-8b / lora / eval_mmlu
[0166] eval
[0167] batch_size: 1
[0168] (7) Large model Q&A application
[0169] Build a local knowledge base. By inputting local pdf files into the system, build a machine learning Q&A knowledge base. For example Figure 11 . Upload a large number of machine learning pdf files, and then through the file processing component, semantically understand, segment, and vectorize the file data and store it in the vector database milvus and the search engine Elasticsearch. Synchronously associate the files with the knowledge base, and store the association information in the mysql database.
[0170] Build a knowledge base, train and infer on the large model Q&A interface. Now the inference result interface is as follows Figure 12 shown.
[0171] Judging from the results of model inference, the effect of Q&A is very good and the accuracy is very high.
[0172] The device provided by the present invention can collect a local knowledge base according to the large model schema, construct large model fine-tuning, train, test, and evaluate the model through the local knowledge base, and realize the generation of accurate Q&A texts in the local knowledge base, thereby reducing the complex and costly model pre-training of the large model and quickly obtaining better model results.
[0173] In this specification, each embodiment is described in a progressive manner. The key point of each embodiment is the difference from other embodiments. For the same and similar parts among the embodiments, reference can be made to each other. For the system disclosed in the embodiment, since it corresponds to the method disclosed in the embodiment, the description is relatively simple. For related parts, reference can be made to the description in the method part.
[0174] In this article, specific examples are used to elaborate on the principle and implementation manner of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention; at the same time, for those of ordinary skill in the art, according to the idea of the present invention, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to the present invention.
Claims
1. A large model LLaMA3 fine-tuned knowledge question-answering dialogue system based on a local knowledge base, characterized by: include: LLaMA3 large model parameter acquisition module, used to obtain LLaMA3 large model parameters; A low-rank decomposition technology model building module is used to build a low-rank decomposition technology model; A model training module, used for training the low-rank decomposition technology to construct a model; A local knowledge base construction module is used to construct a local knowledge base; Training model fine-tuning module, used to fine-tune the training model; Model evaluation module, used to evaluate the model; Model deployment module, used to deploy models.
2. The large model LLaMA3 fine-tuned knowledge question-answering dialogue system based on local knowledge base according to claim 1 is characterized in that: The fine-tuning training model includes: Build a dataset; Build the base model and pre-train the LLaMA basic model; Supervised fine-tuning of SFT for specific tasks; Improve the model's question-answering quality based on human feedback reinforcement learning RLHF; Local knowledge base data enhancement and optimization; Model evaluation, deployment, and continuous optimization.
3. The large model LLaMA3 fine-tuned knowledge question-answering dialogue system based on local knowledge base according to claim 2 is characterized in that: The constructing of the data set specifically includes: constructing the data set using a high-precision model or a large model.
4. The large model LLaMA3 fine-tuning knowledge question-answering dialogue system based on local knowledge base according to claim 2 is characterized in that: The construction of the base model and the pre-training of the LLaMA basic model include: The input feature root mean square normalization RMS Norm is used, the attention mechanism is queried in groups, and residual connections are performed. The sample data is then regularized to ensure that data control produces smaller errors. The residual connection and normal distribution normalization are performed again to complete the entire Decoding calculation.
5. The large model LLaMA3 fine-tuned knowledge question-answering dialogue system based on local knowledge base according to claim 2 is characterized in that: The specific tasks of the supervised fine-tuning SFT are: Supervised fine-tuning is to pre-train a source neural network model on the source dataset, and then use the downstream scene data to create a new target neural network model. The target model copies all model designs and most of its parameters on the source model except the output layer. These model parameters contain the knowledge learned on the source dataset, and this knowledge is also applicable to the target dataset. During fine-tuning, an output layer with an output size equal to the number of categories in the target dataset is added to the target model, and the model parameters of this layer are randomly initialized. When training the target model on the target dataset, it will be trained from scratch to the output layer, and the parameters of the remaining layers are fine-tuned based on the parameters of the source model.
6. The large model LLaMA3 fine-tuning knowledge question-answering dialogue system based on local knowledge base according to claim 2 is characterized in that: The human feedback reinforcement learning RLHF is used to improve the model question answering quality: Supervised fine-tuning uses normal instruction following or conversation samples to train the model's basic conversation and prompt-observing capabilities; based on human preferences and annotations, a scoring model RM that can simulate human preferences is trained; with the help of RM to provide feedback, the model's behavior is continuously adjusted through the PPO reinforcement learning framework.
7. The large model LLaMA3 fine-tuning knowledge question-answering dialogue system based on local knowledge base according to claim 2 is characterized in that: The local knowledge base data enhancement and optimization are specifically as follows: Input prompt word retrieval sequence, embed the query sequence into the word vector model, use the query sequence word vector to match the vector in the vector database for similarity, search the corresponding document in the word vector database, find the top k similarity sequences as the reference sequence vector of the answer, combine the retrieved text sequence with the query sequence, construct a better prompt word prompt, and then pass the prompt word to the large language model. Given the input sequence X, after word vector embedding, X_embeding = embedding(X,V), V is the word vector table, the word vector enters the database for matching, calculates the two to obtain the text X_query = Q(X_embedding,D), and outputs the text as the input series of the large model.
8. The large model LLaMA3 fine-tuned knowledge question-answering dialogue system based on local knowledge base according to claim 2 is characterized in that: The model evaluation, deployment and continuous optimization are specifically as follows: First, download the llama3 model file. Using the pre-trained model + LORA adapter, you can quickly and efficiently adapt the model to new tasks or fields without retraining the entire model. The accuracy of the model after fine-tuning and the sequence evaluation of the model output results; after the local knowledge base is built, the model is trained and the parallel computing of the high-performance GPUA10 architecture of the Aliyun artificial intelligence platform is used to obtain the evaluation indicators of the model; the model evaluation selects a large-scale multi-task language understanding benchmark for testing, and the evaluation method uses questions in each subject area to calculate the accuracy of the model separately, and calculates the average accuracy of all subjects as the overall performance of the model, and compares the performance of the model with the accuracy of human experts.
9. The large model LLaMA3 fine-tuned knowledge question-answering dialogue system based on local knowledge base according to claim 2 is characterized in that: It also includes the establishment of a large-model question-and-answer application platform.
10. The large model LLaMA3 fine-tuned knowledge question-answering dialogue system based on local knowledge base according to claim 9 is characterized in that: The large-model question-and-answer application platform includes: Build a local knowledge base, input local PDF files through the system, build a machine learning question and answer knowledge base, upload a large number of machine learning PDF files, and then use the file processing component to semantically understand, segment, and vectorize the file data and store it in the vector database and search engine. At the same time, associate the file with the knowledge base, store the associated information in the MySQL database, build the knowledge base, and perform training and reasoning in the large model question and answer interface.