Active dialogue model fine tuning method based on knowledge editing, terminal and medium
Through the method of subdividing tasks and efficient parameter module fine-tuning, the executability of the language model in user intention understanding and task execution is improved, the problem of difficult user intentions in the prior art is solved, and more efficient user participation and information interaction are achieved.
Patent Information
- Application Number
- CN202510266535.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-07
- Publication Date
- 2025-07-11
AI Technical Summary
When the existing language model deals with user intention recognition and understanding, especially in open scenarios, it is difficult to effectively explore the real needs of users, and the lack of an effective user participation mechanism, making it difficult to meet the potential intentions of users when performing tasks.
By collecting active dialogue data sets and general pre-trained data sets in specific fields, they are divided into three sub-tasks: fuzzy demand prediction, active questioning and user demand overview, and fine-tuning the language model through parameter efficient modules, using methods such as singular value decomposition and cross-entropy loss function, the model's active questioning ability is improved and user intentions are dynamically mined.
It improves the executability of the language model in user intention understanding and task execution, reduces training time and computing resources, and improves model performance and user experience.
Smart Images

Figure CN120297329A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of dialogue systems, and particularly to an active dialogue model fine-tuning method, a terminal, and a medium based on knowledge editing. Background Art
[0002] In recent years, due to the fact that the LLM (Large Language Model) can process and understand complex language and context information, it has become very popular in performing various natural language tasks and is widely used in dialogue systems, such as intelligent assistants, chatbots, and customer services. How to perform intent recognition through natural language communication between the LLM and users has become the focus of attention. However, in actual scenarios, due to factors such as the ambiguity of instruction expressions and different user focuses, the LLM faces the challenge of "mining and clarifying the user's true needs".
[0003] Traditional demand understanding methods often rely on task-based dialogue systems, aiming to help users complete certain tasks in specific domains, such as restaurant reservations, weather inquiries, and flight bookings, etc. During the dialogue process, intent classification, entity recognition, and slot filling need to be achieved. However, the traditional fixed-template-based demand understanding strategy depends on the design of a single rule and can only solve the understanding of user intents in specific scenarios, facing the problem of being difficult to meet the demand mining in open scenarios.
[0004] Existing work benefits from the rapid development of language models represented by Transformer. People tend to introduce external knowledge bases, such as historical dialogues, domain knowledge, etc., to provide more background knowledge for large models, or stimulate the reasoning ability of large language models through methods such as CoT. Current language model-driven agents usually lack an effective user participation mechanism. Although these agents are good at formulating strategies and executing tasks, they are insufficient in seeking clarification and grasping the accurate user intent. During the execution of agent tasks, although the goals seem to be achieved, it is difficult to meet the user's potential true intent. In addition, with the proposal of the existing active dialogue dataset IN3, it becomes possible for the model to strengthen the interaction with users through active questioning. However, the number of training items in this dataset is too small. Therefore, how to enhance user participation through the interaction between the LLM and users to strengthen information interaction, so that the LLM can further understand user intents, has become an urgent problem to be solved in the current field of dialogue systems. Summary of the Invention
[0005] To solve the technical problems existing in the prior art, the present invention provides an active dialogue model fine-tuning method, a terminal, and a medium based on knowledge editing. The present invention fine-tunes the active dialogue model to realize the active questioning of the model to mine the potential intent of the user, so as to enhance its own understanding of the executability of the task and guide the subsequent task progress.
[0006] To achieve the above object, the present invention provides the following technical solutions:
[0007] The present invention discloses an active dialogue model fine-tuning method based on knowledge editing, comprising the steps of:
[0008] S1. Collect a specific domain active dialogue data set and a general pre-training data set and perform preprocessing. The active dialogue tasks in the data set are divided into three sub-tasks: fuzzy requirement prediction, active question asking, and user requirement overview, and the data format is {instruction, input, output};
[0009] S2. Use the active dialogue data set and the pre-training data set to perform instruction fine-tuning and secondary pre-training on a pre-trained language model added with a parameter-efficient module;
[0010] S3. Extract and enhance specific capabilities on the parameter-efficient module to obtain the adjusted model parameters of the pre-trained language model;
[0011] S4. Based on the adjusted model parameters, trigger the active question asking ability through in-context learning to achieve dynamic mining of user intentions.
[0012] As a further improvement of the above solution, the weight parameters of the parameter-efficient module are represented as W ∈ R M×K , M represents the number of row vectors, and K represents the dimension of the row vectors; among them, the weight of the parameter-efficient module fine-tuned on the general pre-training data set is represented as W o , and the weight of the parameter-efficient module fine-tuned on the active dialogue data set is represented as W + ; R is the set of real numbers.
[0013] As a further improvement of the above solution, step S3 includes the following specific steps:
[0014] S31. Construct an orthogonal vector space of basic capabilities:
[0015] [U, ∑, V T ← svd(W o )
[0016] In the formula, U ∈ R M×M , Σ ∈ R M×K and V ∈ R K×K are respectively the left singular matrix, the singular value matrix, and the right singular matrix obtained from W o through the singular value decomposition algorithm svd(·); the superscript T represents the transpose;
[0017] S32. Feature vector projection:
[0018] ∑ task = U T * (W + - Wo )*V
[0019] Where, ∑ task is the singular value matrix formed after projecting the difference between W + and W o onto the orthogonal vector space formed by U and V;
[0020] S33. Specific ability extraction:
[0021]
[0022] Where, represents the i-th projected singular value matrix, i ∈ [1, M]; By setting a threshold ε to filter out the smaller singular values in to retain relevant features to form a new singular value matrix and the specific ability vector Ext(W), l ∈ [1, M];
[0023] S34. Specific ability enhancement:
[0024]
[0025] Where, W * represents the original model parameters of the pre-trained language model; represents the adjusted model parameters; λ is a hyperparameter used to balance the weights of the simple LoRA fine-tuned model parameters and the specific ability parameters.
[0026] As a further improvement of the above solution, in step S2, the pre-trained language model added with the parameter-efficient module is trained by the cross-entropy loss function and the Adam optimizer.
[0027] As a further improvement of the above solution,
[0028] In step S2, the fine-tuning method uses LoRA fine-tuning, expressed as:
[0029] h = W'x + ΔWx = W'x + BAx
[0030] Where, h are the model parameters of the pre-trained language model after instruction fine-tuning; W' ∈ R d×k represents the pre-trained weight matrix, which is frozen during training; ΔW represents the model parameters introduced during fine-tuning; B ∈ R d×r and A ∈ R r×k are the introduced low-rank matrices; x ∈ R k represents the input hidden state; R is the set of real numbers.
[0031] As a further improvement of the above solution, the expression formula of the cross-entropy loss function is:
[0032]
[0033] In the formula, Loss represents the loss function for model training on the dataset; NUM represents the number of samples, where num ∈ [1, NUM]; S represents the sequence length of the standard output, where s ∈ [1, S]; represents the true label y s at position s
[0034] As a further improvement of the above solution, in step S4, the active questioning ability is triggered by designing prompt words to guide the model to judge the ambiguity of the user's intention and generate clarification questions.
[0035] The present invention also discloses a computer terminal, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of the active dialogue model fine-tuning method based on knowledge editing as described above are implemented.
[0036] The present invention also discloses a computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, the steps of the active dialogue model fine-tuning method based on knowledge editing as described above are implemented.
[0037] Compared with the prior art, the beneficial effects of the present invention are:
[0038] 1. The active dialogue model fine-tuning method based on knowledge editing disclosed by the present invention focuses on the active dialogue scenario of the model, divides the active dialogue task into three subtasks: fuzzy demand prediction, active questioning, and user demand overview, and combines the three subtasks into a standard output in the form of a dialogue, and fine-tunes the language model in an end-to-end manner to improve the overall performance of the model and the user experience.
[0039] 2. Limited by the shortage of training resources, in order to reduce the training time and computing resources, the fine-tuning method of the present invention does not require full-parameter fine-tuning, and adopts the existing method of selectively updating certain weights in the model by adding additional parameters, which can speed up the training process of new tasks while retaining most of the pre-trained knowledge.
[0040] 3. In order to better improve the ability, the fine-tuning method of the present invention designs a model fine-tuning method based on knowledge editing, extracts vector parameters that can achieve specific functions through orthogonal decomposition, projection, and arithmetic operations in the weight space, effectively alleviates the problem of insufficient model training caused by low-rank fine-tuning in the low-resource background, and thus effectively improves the performance of the model to a certain extent.
[0041] 4. The computer terminal and computer-readable storage medium disclosed in the present invention can achieve the same beneficial effects as the above method by applying the above method, and will not be elaborated here. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] Figure 1 It is a flowchart of the method for fine-tuning an active dialogue model based on knowledge editing in Embodiment 1 of the present invention.
[0043] Figure 2 It is a framework diagram of the method for fine-tuning an active dialogue model based on knowledge editing in Embodiment 1 of the present invention.
[0044] Figure 3 It is a schematic structural diagram of a computer terminal in Embodiment 2 of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0045] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the protection scope of the present invention.
[0046] Embodiment 1
[0047] Please refer to Figures 1 to 2 , this embodiment provides a method for fine-tuning an active dialogue model based on knowledge editing, including steps S1 to S4.
[0048] S1. Collect an active dialogue data set in a specific domain and perform preprocessing. The active dialogue tasks in the data set are divided into three subtasks: fuzzy requirement prediction, active question asking, and user requirement overview, and the data format is {instruction, input, output}.
[0049] Collect a general pre-training data set and perform preprocessing, and select some subsets to accelerate the fine-tuning process.
[0050] In this embodiment, two data sets publicly available on the network are selected: the active dialogue data set IN3 and the general pre-training data set C4.
[0051] The present invention divides the active dialogue model into three subtasks. There is a logical relationship between the three subtasks and similar capabilities are required to a certain extent to improve the effect of the model in real-world applications.
[0052] S2. Use the active dialogue data set to perform instruction fine-tuning and secondary pre-training on a pre-trained language model added with a parameter-efficient module.
[0053] S2. Use the active dialogue dataset and the pre-training dataset to perform instruction fine-tuning and secondary pre-training on the pre-trained language model with a parameter-efficient module added.
[0054] In this embodiment, step S2 may include the following steps:
[0055] S21. Select the base pre-trained language model, the fine-tuning method, and the parameters. The fine-tuning method for the active dialogue model is as follows:
[0056] Use the template recommended by the pre-trained language model to combine instruction i , x i and y i . Here, the parameter fine-tuning method is exemplified by LoRA fine-tuning. Select the LoRA parameters of the same content (such as: rank, specified model position, etc.) for subsequent combination of model parameters. W ∈ R M×K ; R is the set of real numbers; where M represents the number of row vectors and K represents the dimension of the row vectors. The LoRA fine-tuning is expressed as:
[0057] h = W'x + ΔWx = W'x + BAx
[0058] In the formula, h is the model parameter of the pre-trained language model after instruction fine-tuning; W' ∈ R d×k represents the pre-trained weight matrix, which is frozen during the training process; ΔW represents the model parameter introduced during the fine-tuning process; B ∈ R d×r and A ∈ R r×k are the introduced low-rank matrices; x ∈ R k represents the input hidden state; R is the set of real numbers.
[0059] S22. Train the model by predicting the next most likely character through the large language model, and construct the cross-entropy loss function Loss:
[0060]
[0061] In the formula, Loss represents the loss function for model training on the dataset (the same formula is used for fine-tuning on both datasets); NUM represents the number of samples, num ∈ [1, NUM]; S represents the length of the standard output sequence, s ∈ [1, S]; represents the softmax probability of the true label y s at position s.
[0062] S23. Use the Adam optimizer to train these two models separately on 2 NVIDIA GeForce RTX 4090s, and minimize the total loss function Loss to update the model parameters. End after training for 5 epochs.
[0063] S3. Extract and enhance specific capabilities on the parameter-efficient module to obtain the adjusted model parameters of the pre-trained language model.
[0064] Step S3 includes the following specific steps:
[0065] S31. Construct an orthogonal vector space of basic capabilities:
[0066] [U,∑,V T ←svd(W o )
[0067] Where U ∈ R M×M , Σ ∈ R M×K and V ∈ R K×K are the left singular matrix, singular value matrix, and right singular matrix obtained from W o by the singular value decomposition algorithm svd(·), respectively; the superscript T represents the transpose;
[0068] S32. Feature vector projection:
[0069] Σ task =U T *(W + -W o )*V
[0070] Where Σ task is the singular value matrix formed by projecting the difference between W + and W o onto the orthogonal vector space formed by U and V;
[0071] S33. Specific ability extraction:
[0072]
[0073] Where represents the i-th projected singular value matrix, i ∈ [1, M]; by setting a threshold ε to filter out the smaller singular values in it, retaining the relevant features to form a new singular value matrix and the specific ability vector Ext(W), l ∈ [1, M];
[0074] S34. Specific ability enhancement:
[0075]
[0076] Where W * represents the original model parameters of the pre-trained language model; represents the adjusted model parameters; λ is a hyperparameter used to balance the weights of the simple Lora fine-tuned model parameters and the specific ability parameters.
[0077] The present invention reduces the model training time and computing resources through simple projection and linear operations. Additionally, by reducing the requirements for the current task and relying instead on other general pre-trained datasets, the specific capabilities of large models are enhanced.
[0078] S4. Based on the adjusted model parameters, trigger the active questioning ability through in-context learning to dynamically mine the user's intentions.
[0079] In step S4, the active questioning ability is triggered by designing prompt words to guide the model to judge the ambiguity of the user's intention and generate clarification questions.
[0080] To verify the effectiveness of the method of the present invention, the only publicly available dataset IN3 in active dialogue research is adopted in this embodiment. Training is carried out on the training set and verification is carried out on the validation set.
[0081] In the active dialogue task, this embodiment considers using Mistral-7B-Instruct-v0.2 as the base model and performing LoRA fine-tuning on the basis of the base model, with the trainable parameters being 0.2888%. And, to enable automated construction, the Mistral-Nemo-12b model is deployed locally through ollama as an agent model to implement autonomous question-answering on behalf of the user. For each task, this embodiment prompts the model to clearly judge the task ambiguity, ask for missing details, and summarize the user's goal. By recording the dialogue process, and then evaluating through direct statistical calculation and the agent model Mistral-7B-Instruct-v0.2 according to the standard answers provided in IN3.
[0082] To verify the executability and generality of the method, we additionally apply the method to the analogical reasoning task, and we still use Mistral-7B-Instruct-v0.2 for the experiment.
[0083] In the active dialogue task, according to IN3, the model is evaluated mainly from three aspects.
[0084] (1) ACC Ambiguity Judgment Accuracy: Calculate the percentage of the model's judgment of the ambiguity (ambiguous or clear) of task t that is consistent with the ground truth. This measures the model's ability to recognize ambiguity and clarity and avoid asking questions about tasks that are already clear.
[0085] (2) Recover iAverage missing detail recovery rate: In the IN3 dataset, there are three options with different importance indicators for each question. The larger the number, the more useful the detail is for restoring the user's true intention. For the truly missing details of different importance levels, the analysis model in this embodiment analyzes the average number of dialogue turns restored (explicit queries) during the interaction process. This measures the model's ability to ask necessary details in the shortest possible time.
[0086] (3)、efficiency i Efficiency: To enhance the user experience, we additionally introduce an indicator to measure the ratio between ACC and Recover i between them.
[0087] In the analogical reasoning task, this embodiment mainly considers the following two aspects for experiments.
[0088] (1)、Inference accuracy: In multiple-choice questions, points are awarded according to the final answer, with points for correct answers and no points for incorrect answers.
[0089] (2)、Explanation generation: Only the question and options are provided, and the LLM needs to give the solution ideas for the question and options respectively, and then generate the final option.
[0090] The active dialogue in this embodiment selects a simple Lora fine-tuning model for effect comparison. Specifically, Table 1 shows the experimental results in the above dataset. It can be observed that there is still room for improvement in the simple Lora fine-tuning. In contrast, the method of the present invention improves the overall performance by selecting specific hyperparameters while ensuring the basic capabilities of the model are maintained.
[0091] Table 1: Active dialogue experiment: Results
[0092]
[0093]
[0094] The analogical reasoning in this embodiment selects a simple Lora fine-tuning model and the existing research Uniform-soup for effect comparison. Uniform-soup is: Model soups: averaging weights of multiple fine-tuned models improves accuracy without increasing inference time.
[0095] Table 2: Category reasoning experimental results
[0096]
[0097] As can be seen from Table 2, the method of the present invention is superior to existing research in analogical reasoning. The method of Uniform-soup does not necessarily produce better performance. By setting different hyperparameters, the accuracy of the method of the present invention can be further improved.
[0098] Example 2
[0099] This embodiment provides a computer terminal, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of the active dialogue model fine-tuning method based on knowledge editing as described in Example 1 are implemented.
[0100] As Figure 3 shown, the computer terminal provided in this embodiment includes: at least one processor 101, and a memory 102 connected to at least one processor 101. In this embodiment, the specific connection medium between the processor 101 and the memory 102 is not limited. Figure 3 Here, it is taken as an example that the processor 101 and the memory 102 are connected through a bus 100. The bus 100 is represented by a thick line in Figure 3 Here. The connection manners between other components are only illustrative and not limiting. The bus 100 can be divided into an address bus, a data bus, a control bus, etc. For the convenience of representation, Figure 3 it is only represented by a thick line in here, but it does not mean that there is only one bus or one type of bus. Alternatively, the processor 101 can also be called a controller, and there is no limitation on the name.
[0101] In this embodiment, the memory 102 stores instructions executable by at least one processor 101. By executing the instructions stored in the memory 102, at least one processor 101 can execute the foregoing method.
[0102] Among them, the processor 101 is the control center of the terminal, and can connect various parts of the entire control device through various interfaces and lines. By running or executing the instructions stored in the memory 102 and calling the data stored in the memory 102, various functions of the terminal and process data, so as to monitor the terminal as a whole.
[0103] In a possible design, the processor 101 may include one or more processing units. The processor 101 may integrate an application processor and a modulation and demodulation processor. Among them, the application processor mainly processes the operating system, user interface, application programs, etc., and the modulation and demodulation processor mainly processes wireless communication. It can be understood that the above modulation and demodulation processor may not be integrated into the processor 101. In some embodiments, the processor 101 and the memory 102 can be implemented on the same chip, and in some embodiments, they can also be separately implemented on independent chips.
[0104] The processor 101 may be a general-purpose processor, such as a central processing unit (CPU), a digital signal processor, an application specific integrated circuit, a field programmable gate array, or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, and can implement or execute the various methods, steps, and logic block diagrams disclosed in this embodiment. The general-purpose processor may be a microprocessor or any conventional processor, etc. The steps of the active dialogue model fine-tuning method based on knowledge editing disclosed in conjunction with Embodiment 1 can be directly embodied as being executed by the hardware processor, or executed by a combination of hardware and software modules in the processor 101.
[0105] As a non-volatile computer-readable storage medium, the memory 102 can be used to store non-volatile software programs, non-volatile computer-executable programs, and modules. The memory 102 may include at least one type of storage medium, for example, it may include flash memory, hard disk, multimedia card, card-type memory, random access memory (RAM), static random access memory (SRAM), programmable read-only memory (PROM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), magnetic memory, magnetic disk, optical disc, etc. The memory 102 is any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto. The memory 102 in this embodiment may also be a circuit or any other device capable of implementing a storage function, for storing program instructions and / or data.
[0106] By designing and programming the processor 101, the code corresponding to the active dialogue model fine-tuning method based on knowledge editing introduced in the foregoing embodiment can be solidified into the chip, so that the chip can execute Figure 1 the steps of the active dialogue model fine-tuning method based on knowledge editing as shown. How to design and program the processor 101 is well-known to those skilled in the art and will not be elaborated herein.
[0107] Embodiment 3
[0108] This embodiment provides a computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, the steps of the active dialogue model fine-tuning method based on knowledge editing as described in Embodiment 1 are implemented.
[0109] The computer-readable storage medium may include flash memory, a hard disk, a multimedia card, a card-type memory (such as an SD or DX memory, etc.), a random access memory (RAM), a static random access memory (SRAM), a read-only memory (ROM), an electrically erasable programmable read-only memory (EEPROM), a programmable read-only memory (PROM), a magnetic memory, a magnetic disk, an optical disk, etc. In some embodiments, the storage medium may be an internal storage unit of a computer device, such as the hard disk or memory of the computer device. In other embodiments, the storage medium may also be an external storage device of the computer device, such as a plug-in hard disk equipped on the computer device, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc. Of course, the storage medium may also include both the internal storage unit and the external storage device of the computer device. In this embodiment, the memory is generally used to store the operating system and various application software installed on the computer device. In addition, the memory can also be used to temporarily store various data that have been output or will be output.
[0110] As mentioned above, the above are only the preferred specific embodiments of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention, according to the technical solution and inventive concept of the present invention, makes equivalent substitutions or changes, and should be covered by the protection scope of the present invention.
Claims
1. A fine-tuning method for an active dialogue model based on knowledge editing, characterized in that Including the steps: S1. Collect an active dialogue dataset in a specific domain and a general pre-training dataset and perform preprocessing. The active dialogue tasks in the dataset are divided into three subtasks: fuzzy requirement prediction, active question asking, and user requirement overview, and the data format is {instruction, input, output}; S2. Use the active dialogue dataset and the pre-training dataset to perform instruction fine-tuning and secondary pre-training on a pre-trained language model with a parameter-efficient module added; S3. Extract and strengthen specific capabilities on the parameter-efficient module to obtain the adjusted model parameters of the pre-trained language model; S4. Based on the adjusted model parameters, trigger the active question-asking ability through in-context learning to achieve dynamic mining of user intentions.
2. The fine-tuning method of an active dialogue model based on knowledge editing according to claim 1, wherein The weight parameters of the parameter-efficient module are represented as W ∈ R M×K , where M represents the number of row vectors and K represents the dimension of the row vectors; among them, the weights of the parameter-efficient module fine-tuned on the general pre-training dataset are represented as W o , and the weights of the parameter-efficient module fine-tuned on the active dialogue dataset are represented as W + ; R is the set of real numbers.
3. The fine-tuning method of an active dialogue model based on knowledge editing according to claim 2, characterized in that, Step S3 includes the following specific steps: S31. Construct an orthogonal vector space for basic capabilities: [U,∑,V T ←svd(W o ) where \(U\in\mathbb{R}\) M×M ,\(\Sigma\in\mathbb{R}\) M×K and \(V\in\mathbb{R}\) K×K are the left singular matrix, the singular value matrix, and the right singular matrix obtained from \(W\) o by the singular value decomposition algorithm \(\text{svd}(\cdot)\); the superscript \(T\) represents the transpose; S32. Feature vector projection: Σ task = U T *(W + - W o )*V where, ∑ task is the singular value matrix formed after projecting the difference between W + and W o onto the orthogonal vector space formed by U and V; S33. Specific capability extraction: Wherein, represents the i-th projected singular value matrix, i ∈ [1, M]; By setting a threshold ε to filter out the smaller singular values in, and retain relevant features to form a new singular value matrix and a specific ability vector Ext(W), l ∈ [1, M]; S34. Specific capability strengthening: Where W * represents the original model parameters of the pre-trained language model; represents the adjusted model parameters; λ is a hyperparameter used to balance the weights of the simple LoRA fine-tuned model parameters and the specific ability parameters.
4. A method for fine-tuning an active dialogue model based on knowledge editing according to claim 1, characterized in that, In step S2, the pre-trained language model with the parameter-efficient module added is trained through a cross-entropy loss function and an Adam optimizer.
5. A method for fine-tuning an active dialogue model based on knowledge editing according to claim 4, characterized in that, In step S2, the fine-tuning method uses LoRA fine-tuning, expressed as: h = W′x + ΔWx = W′x + BAx where h are the model parameters of the pre-trained language model fine-tuned by an instruction; W′ ∈ R d×k represents the pre-trained weight matrix, which is frozen during the training process; ΔW represents the model parameters introduced during the fine-tuning process; B ∈ R d×r and A ∈ R r×k are the introduced low-rank matrices; x ∈ R k represents the input hidden state; R is the set of real numbers.
6. A method for fine-tuning an active dialogue model based on knowledge editing according to claim 5, characterized in that, The expression formula of the cross-entropy loss function is: Where Loss represents the loss function for model training on the dataset; NUM represents the number of samples, num ∈ [1, NUM]; S represents the length of the sequence of the standard output, s ∈ [1, S]; represents the true label y at position s s of the softmax probability.
7. A fine-tuning method for an active dialogue model based on knowledge editing according to claim 1, characterized in that, In step S4, the active question-asking ability is triggered by designing prompt words to guide the model to discriminate the ambiguity of the user's intention and generate clarification questions.
8. A computer terminal, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, the steps of the active dialogue model fine-tuning method based on knowledge editing according to any one of claims 1 to 7 are implemented.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, the steps of the active dialogue model fine-tuning method based on knowledge editing according to any one of claims 1 to 7 are implemented.
Citation Information
Cited By
AI voice interaction-based pension service scheduling method and system
CN120748370A