Dialogue abstract automatic generation method and system based on fine tuning large language model
By using different prompt word pairs in the fine-tuning training and inference stage of large language model, the problem that the model is difficult to balance simplicity and accuracy in the dialogue summary task is solved, and a more efficient dialogue summary generation is achieved.
Patent Information
- Application Number
- CN202510222222.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-27
- Publication Date
- 2025-06-20
AI Technical Summary
Existing large language models are difficult to balance the simplicity and precision of generating abstracts in conversation summary tasks, and often generate detailed abstracts without simplicity or short summary loses key information.
By designing different prompt words to form prompt word pairs during the model's fine-tuning training and inference verification stage, using strict instructions during the training stage and simple instructions during the reasoning stage to balance the simplicity and accuracy of the dialogue summary output.
It improves the simplicity and accuracy of the model when generating dialogue summary, helps the model better understand the dialogue content, reduce noise, and improves the robustness in NLP tasks.
Smart Images

Figure CN120179815A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of natural language deep learning processing, and particularly to a method and system for automatically generating dialogue summaries based on fine-tuning large language models. Background Art
[0002] With the popularization of the Internet and instant messaging platforms, the communication between people has become closer, and dialogue texts have become a common data form. Therefore, as a relatively popular research direction in the text summarization task - dialogue summarization has emerged, aiming to extract and summarize the core information from the original dialogue. In the era of LLM, dialogue summarization has been regarded as an important indicator to measure the ability of LLM. The summaries written by existing LLMs already conform to human habits, but there are sometimes problems in dialogue-type texts, mainly because after the LLM obtains this type of dialogue input, due to its inherent generation mechanism, it tends to continue generating coherent dialogues in the previous format rather than summarizing.
[0003] With the rise of LLM, how to effectively utilize the powerful capabilities of these models has become a research hotspot. After continuous exploration by scholars, the academic community generally believes that efficient parameter fine-tuning training (PEFT) of open-source LLMs can improve the performance in NLP tasks without large-scale parameter retraining. The performance of LLM in the dialogue summarization task has the following deficiencies: The open-source LLM without fine-tuning training usually generates detailed summaries lacking conciseness and tends to continue generating coherent dialogues in the previous format rather than summarizing. The fine-tuning method can improve the model's ability to summarize the summary to a certain extent by imitating the "input / output" format of the training data, but the key information in the dialogue is scattered, and the fine-tuning method is prone to causing the LLM to miss important information and the score to drop. The LLM is in a dilemma of being difficult to find a balance between the conciseness and accuracy of the summary. Summary of the Invention
[0004] Therefore, the present invention provides a method and system for automatically generating dialogue summaries based on fine-tuning large language models, which solves the problem that it is difficult for existing large models to balance the conciseness and accuracy of generated summaries, uses different prompt words to form prompt word pairs, uses strict instructions in the training stage, and simple instructions in the inference stage, and balances the conciseness and accuracy of dialogue summary output through different prompt words.
[0005] According to the design scheme provided by the present invention, on the one hand, a method for automatically generating dialogue summaries based on fine-tuning large language models is provided, including:
[0006] Set a first summary instruction and obtain a training dialogue summary corpus dataset. The first summary instruction is a prompt instruction label indicating the generation of a dialogue summary with a fixed length. The training dialogue summary corpus dataset contains a number of dialogue data samples and reference dialogue summaries. Each dialogue data sample is dialogue data information of consensus or decision-making communication carried out by at least one dialogue role around a specific dialogue task;
[0007] Divide the training dialogue summary corpus dataset into a dialogue summary training dataset and a dialogue summary test dataset according to a preset ratio. Combine the first summary instruction with the dialogue summary training dataset to form a first dialogue summary template, and use the first dialogue summary template to fine-tune and train an open-source large language model to obtain a fixed-length summary generation model;
[0008] Set a second summary instruction. The second summary instruction is a prompt instruction label indicating the generation of a dialogue summary without a length limit. Combine the second summary instruction with the dialogue summary test dataset to form a second dialogue summary template, use the second dialogue summary template to evaluate and optimize the fixed-length summary generation model, and obtain a dialogue summary target generation model based on the evaluation and optimization results;
[0009] Combine the target dialogue corpus data with the second summary instruction to form a second dialogue summary template, and use the dialogue summary target generation model to generate a dialogue summary of the target corpus data. The target dialogue corpus data is the dialogue data for which a dialogue summary is to be generated.
[0010] As the dialogue summary automatic generation method based on fine-tuning a large language model of the present invention, further, obtaining a training dialogue summary corpus dataset includes:
[0011] Obtain a publicly available dialogue summary dataset. The publicly available dialogue summary dataset includes: a dialogue dataset constructed by translating daily monolingual dialogues from a source language to a target language, and a dialogue dataset constructed by collecting cross-lingual daily dialogue pairs;
[0012] Select a corresponding dialogue dataset from the publicly available dialogue summary dataset according to the dialogue summary automatic generation task scenario and based on the publicly available dialogue summary dataset, and use the selected dialogue dataset as the training dialogue summary corpus dataset.
[0013] As the dialogue summary automatic generation method based on fine-tuning a large language model of the present invention, further, forming a dialogue summary template includes:
[0014] Analyze the dialogue roles in the dialogue data, format the dialogue data according to the dialogue roles, and fill or insert the dialogue summary as a summary label into the formatted dialogue data.
[0015] As the method for automatically generating conversation summaries based on fine-tuning large language models in the present invention, further, the open-source large language model adopts a Chinese-English bilingual conversation pre-training model or a lightweight pre-training language model.
[0016] As the method for automatically generating conversation summaries based on fine-tuning large language models in the present invention, further, use a second conversation summary template to evaluate and optimize the fixed-length summary generation model, including:
[0017] Evaluate the accuracy and conciseness of the model-generated summary based on the model-generated summary and the reference conversation summary;
[0018] Adjust the first summary instruction and / or the second summary instruction according to the evaluation results to determine the correctness of the set summary instruction.
[0019] As the method for automatically generating conversation summaries based on fine-tuning large language models in the present invention, further, evaluate the accuracy of the model-generated summary, including:
[0020] Divide the reference conversation summary and the model-generated summary into segments according to a sliding window of n words respectively to form n-grams of length n, and calculate the number of n-grams that appear simultaneously in the reference conversation summary and the model-generated summary, and calculate the ROUGE-N metric value for evaluating the accuracy of the model-generated summary based on this number;
[0021] Obtain the longest common subsequence in the reference conversation summary and the model-generated summary, and calculate the ROUGE-L metric value for evaluating the accuracy of the model-generated summary based on the longest common subsequence;
[0022] Perform a weighted calculation on the accuracy of the model-generated summary based on the two metric values of ROUGE-N and ROUGE-L. If the calculation result reaches the accuracy threshold, it is determined that the model-generated summary meets the accuracy expectation standard, and obtain the target generation model for the conversation summary.
[0023] As the method for automatically generating conversation summaries based on fine-tuning large language models in the present invention, further, evaluate the conciseness of the model-generated summary, including:
[0024] Calculate the length penalty factor based on the length of the reference conversation summary and the length of the model-generated summary, and use the length penalty factor and calculate the BLEU metric value for evaluating the conciseness of the model-generated summary based on the number of n-grams in the reference conversation summary and the model-generated summary. If the change in the BLEU metric values of the reference conversation summary and the model-generated summary does not exceed the preset threshold, it is determined that the model-generated summary meets the conciseness expectation standard, and obtain the target generation model for the conversation summary.
[0025] On the other hand, the present invention also provides a dialogue summary automatic generation system based on fine-tuning a large language model, comprising: an instruction generation module, a model fine-tuning module, a model evaluation module, and a summary generation module, wherein,
[0026] The instruction generation module is used to set a first summary instruction and obtain a training dialogue summary corpus dataset. The first summary instruction is a prompt instruction label indicating the generation of a dialogue summary with a fixed length. The training dialogue summary corpus dataset includes a number of dialogue data samples and reference dialogue summaries. Each dialogue data sample is dialogue data information of consensus or decision-making communication carried out by no less than one dialogue role around a specific dialogue task;
[0027] The model fine-tuning module is used to divide the training dialogue summary corpus dataset into a dialogue summary training dataset and a dialogue summary test dataset according to a preset ratio, form a first dialogue summary template with the first summary instruction and the dialogue summary training dataset, and use the first dialogue summary template to fine-tune and train an open-source large language model to obtain a fixed-length summary generation model;
[0028] The model evaluation module is used to set a second summary instruction. The second summary instruction is a prompt instruction label indicating the generation of a dialogue summary without a length limit, and form a second dialogue summary template with the second summary instruction and the dialogue summary test dataset, use the second dialogue summary template to evaluate and optimize the fixed-length summary generation model, and obtain a dialogue summary target generation model according to the evaluation and optimization results;
[0029] The summary generation module is used to form a second dialogue summary template with the target dialogue corpus data and the second summary instruction, and use the dialogue summary target generation model to generate a dialogue summary of the target corpus data. The target dialogue corpus data is the dialogue data for which a dialogue summary is to be generated.
[0030] Advantages of the present invention:
[0031] In the fine-tuning training and inference verification stages of the large model of the present invention, different prompt words are designed to form prompt word pairs, providing clear labels and guidance in the training set, condensing the information of the dialogue summary in a limited space, helping the model to better understand the dialogue content and reduce noise, improving the ability to understand important information, and ensuring the training quality of the LLM; in the inference stage, the simplified instructions give the model greater freedom, enabling it to generate results autonomously according to the content of the context and the task patterns learned during training, enhancing the simplicity of the LLM during summary generation, reducing redundancy, enabling the model to find a balance between the simplicity and accuracy of generating dialogue summaries, and enhancing the robustness of the large model in NPL tasks. Description of the Drawings
[0032] Figure 1Schematic diagram of the automatic dialogue summary generation process based on fine-tuning large language models in the embodiment;
[0033] Figure 2 Schematic diagram of the comparison of summary generation methods in the embodiment;
[0034] Figure 3 Schematic diagram of the principle of the automatic dialogue summary generation algorithm in the embodiment;
[0035] Figure 4 Schematic diagram of the summary scores generated after fine-tuning and training the model with different prompt words in the embodiment;
[0036] Figure 5 Schematic diagram of the experimental hyperparameters in the automatic dialogue summary generation algorithm in the embodiment;
[0037] Figure 6 Schematic diagram of the experimental results of major models in the corresponding datasets and fine-tuning schemes in the embodiment;
[0038] Figure 7 Example of a single dialogue data sample in the dataset in the embodiment. Detailed implementation manner
[0039] To make the objectives, technical solutions, and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below with reference to the accompanying drawings and technical solutions.
[0040] Large language models have the following deficiencies in the dialogue summary task: The zero-shot summaries of open-source large models without fine-tuning training usually generate detailed summaries lacking conciseness, and the large models after fine-tuning usually generate short summaries missing key information. To balance conciseness and accuracy, the embodiment of the present invention, as shown in Figure 1 , provides an automatic dialogue summary generation method based on fine-tuning large language models, including:
[0041] S101. Set a first summary instruction and obtain a training dialogue summary corpus dataset. The first summary instruction is a prompt instruction label indicating the generation of a fixed-length dialogue summary. The training dialogue summary corpus dataset includes several dialogue data samples and reference dialogue summaries. Each dialogue data sample is dialogue data information of at least one dialogue role carrying out consensus or decision-making communication around a specific dialogue task.
[0042] Specifically, obtaining the training dialogue summary corpus dataset can be designed to include:
[0043] Obtain a publicly available dialogue summary dataset, where the publicly available dialogue summary dataset includes: a dialogue dataset constructed by translating daily monolingual conversations from the source language to the target language, and a dialogue dataset constructed by collecting cross-language daily dialogue pairs;
[0044] Automatically generate a task scenario based on the dialogue summary, and select the corresponding dialogue dataset from the publicly available dialogue summary dataset. Use the selected dialogue dataset as the training dialogue summary corpus dataset.
[0045] In the embodiments of this case, the publicly available dialogue summary dataset may include the DialogSum dataset and the SAMSum dataset. Among them, the dialogue samples in the SAMSum dataset are daily life dialogues on a social chat program created by linguists, including many slang words, emojis, and even spelling mistakes; the DialogSum dataset is a dialogue of daily chitchat, which pays more attention to real scenarios.
[0046] S102. Divide the training dialogue summary corpus dataset into a dialogue summary training dataset and a dialogue summary test dataset according to a preset ratio. Combine the first summary instruction with the dialogue summary training dataset to form a first dialogue summary template, and use the first dialogue summary template to fine-tune and train the open-source large language model to obtain a fixed-length summary generation model.
[0047] The open-source large model can adopt the ChatGLM3-6B model and the Gemma-7B-it model. Among them, ChatGLM3 is an open-source Chinese-English bilingual pre-trained model jointly released by Zhipu AI and the KEG Laboratory of Tsinghua University, and Gemma is a lightweight pre-trained language model open-sourced by Google AI. As Figure 2 shown, the zero-shot summary of the open-source large model without fine-tuning training usually generates a detailed summary that lacks conciseness, and the fine-tuned large model usually generates a short summary that loses key information; in the embodiments of this case, in order to balance conciseness and accuracy, different prompt words composed of the first summary instruction and the second summary instruction are set in the training and inference stages of the large model fine-tuning to guide the model to follow the instructions during training and make reasonable play during inference, improving the summary performance.
[0048] Among them, when forming the dialogue summary template, by parsing the dialogue roles in the dialogue data and formatting the dialogue data according to the dialogue roles, the dialogue summary is filled or inserted into the formatted dialogue data as a summary label.
[0049] In order to enhance during training, the language model needs to comply with strict instructions, strengthen the understanding of the training data, and generate short summaries. As Figure 3 shown, the prompt words for this stage can be set to "summarize the following dialogue in one sentence:".
[0050] The prompt words with different length limits are compared. Figure 4It can be seen that better results are obtained with more stringent restrictions.
[0051] Before training the LLM, the data needs to be formatted, assembling the prompts and training data into dialogue-formatted data. Different LLMs support different dialogue templates. When using the ChatGLM3 model as the base model, the template is as follows:
[0052] <|user|>{{P\nX_train\n}}<|assistant|>{{Y_train}}
[0053] When using Gemma-7B as the base model, the template is as follows:
[0054] <start_of_turn>user\n{{P\nX_train\n}}<end_of_turn><start_of_turn>model\n{{Y_train}}<end_of_turn>
[0055] Where P is the prompt, X_train is the original text in the summary dataset training data, and Y_train is the corresponding summary. \n represents the line break character.
[0056] Due to the different lengths of the dataset samples, when training on the DialogSum dataset, the input length can be set to 320 tokens, and when training on the SAMSum dataset, the input length can be set to 256 tokens; the output length is 128 tokens.
[0057] In the fine-tuning of large model training, the P-tuning V2 method and the LoRA method can be used to pre-train the model without changing the original LLM parameters.
[0058] For the P-Tuning V2 method, a trainable prefix sequence length can be set, and this length is usually set to 128.
[0059] Only adjust the weights of an additional part of the parameters during training. Taking the P-tuning V2 method as an example, the total number of parameters optimized during training is as follows:
[0060] Number of parameters = Prefix Sequence Length × Key Value Channels × 4
[0061] For the LoRA method, the low-rank target matrices can be set as the K matrix and the Q matrix, the scaling parameter (lora_alpha) is set to 16, the rank size is 8, and the dropout probability is set to 0.1.
[0062] S103. Set a second summary instruction, where the second summary instruction is a prompt instruction label indicating the generation of a dialogue summary without length limit, and form a second dialogue summary template by combining the second summary instruction with the dialogue summary test data set. Use the second dialogue summary template to evaluate and optimize the fixed-length summary generation model, and obtain the dialogue summary target generation model according to the evaluation and optimization results.
[0063] Before inferring the LLM, it is also necessary to format the data, but at this time, only the prompt words and the original text of the data samples need to be assembled, and the LLM is used to predict the output. Hyperparameters such as Figure 5 As shown, by adjusting appropriate training hyperparameters, the training loss function can be converged; by adjusting different summary task instructions, the correctness of the selected instructions can be determined through inference and evaluation.
[0064] Specifically, using the second dialogue summary template to evaluate and optimize the fixed-length summary generation model includes:
[0065] Evaluating the accuracy and conciseness of the model-generated summary based on the model-generated summary and the reference dialogue summary;
[0066] Adjust the first summary instruction and / or the second summary instruction according to the evaluation results to determine the correctness of the set summary instruction.
[0067] Among them, evaluating the accuracy of the model-generated summary may include:
[0068] Divide the reference dialogue summary and the model-generated summary into segments according to a sliding window of n words respectively to form n-grams of length n, and calculate the number of n-grams that appear simultaneously in the reference dialogue summary and the model-generated summary, and calculate the ROUGE-N metric value for evaluating the accuracy of the model-generated summary based on this number;
[0069] Obtain the longest common subsequence in the reference dialogue summary and the model-generated summary, and calculate the ROUGE-L metric value for evaluating the accuracy of the model-generated summary based on the longest common subsequence;
[0070] Perform a weighted calculation on the accuracy of the model-generated summary based on the two metric values of ROUGE-N and ROUGE-L. If the calculation result reaches the accuracy threshold, it is determined that the model-generated summary meets the accuracy expectation standard, and the dialogue summary target generation model is obtained.
[0071] For the accuracy evaluation of the generated summary, the methods of ROUGE-N (N = 1, 2) and ROUGE-L are used for automatic evaluation. The formula of ROUGE-N is as follows:
[0072]
[0073] Among them, n-gram represents the n-gram, {Ref} represents the reference abstract, and Count match represents the number of occurrences that appear simultaneously in the abstract result and the reference abstract.
[0074] The formula for ROUGE-L is as follows:
[0075]
[0076] Among them, LCS represents the longest common subsequence, and b is a tuning factor used to balance the relative importance of precision and recall.
[0077] Among them, to evaluate the conciseness of the abstract generated by the model, it can include:
[0078] Calculate the length penalty factor based on the length of the reference dialogue abstract and the length of the abstract generated by the model. Use the length penalty factor and calculate the BLEU metric value for evaluating the conciseness of the abstract generated by the model based on the number of n-grams in the reference dialogue abstract and the abstract generated by the model. If the change in the BLEU metric values of the reference dialogue abstract and the abstract generated by the model does not exceed the preset threshold, it is determined that the abstract generated by the model meets the expected standard of conciseness, and the target generation model for the dialogue abstract is obtained.
[0079] For the evaluation of the conciseness of the generated abstract, the results produced by the BLEU and ROUGE methods can be compared. The calculation method of BLEU is as follows:
[0080]
[0081] Among them, w n represents the weight of the n-gram, BP is the length penalty factor, lc is the length of the reference abstract, and lr is the length of the automatic abstract. When the length of the automatic abstract exceeds that of the reference abstract by a large amount, the BLEU score will drop significantly.
[0082] The results of the automatic evaluation method are as Figure 6 shown, and the results of the sample comparison are as Figure 7 shown. Through the sample results, it can be further verified that the solution of this case can improve the abstract performance by setting different prompt word pairs in the training and inference stages of the large model fine-tuning to guide the model to follow the instructions during training and perform reasonably during inference.
[0083] S104. Combine the target dialogue corpus data with the second abstract instruction to form a second dialogue abstract template, and use the target generation model for the dialogue abstract to generate the dialogue abstract of the target corpus data, where the target dialogue corpus data is the dialogue data for which the dialogue abstract is to be generated.
[0084] Deploy the dialogue summary target generation model to the dialogue summary task application scenario to generate dialogue summary data that meets the expectations of conciseness and accuracy.
[0085] Furthermore, based on the above method, an embodiment of the present invention also provides an automatic dialogue summary generation system based on fine-tuning a large language model, including: an instruction generation module, a model fine-tuning module, a model evaluation module, and a summary generation module, where,
[0086] The instruction generation module is used to set a first summary instruction and obtain a training dialogue summary corpus dataset. The first summary instruction is a prompt instruction label indicating the generation of a dialogue summary with a fixed length. The training dialogue summary corpus dataset includes several dialogue data samples and reference dialogue summaries. Each dialogue data sample is dialogue data information of a consensus or decision-making exchange among at least one dialogue role around a specific dialogue task;
[0087] The model fine-tuning module is used to divide the training dialogue summary corpus dataset into a dialogue summary training dataset and a dialogue summary test dataset according to a preset ratio, form a first dialogue summary template with the first summary instruction and the dialogue summary training dataset, and use the first dialogue summary template to fine-tune and train an open-source large language model to obtain a fixed-length summary generation model;
[0088] The model evaluation module is used to set a second summary instruction. The second summary instruction is a prompt instruction label indicating the generation of a dialogue summary without a length limit, and form a second dialogue summary template with the second summary instruction and the dialogue summary test dataset, use the second dialogue summary template to evaluate and optimize the fixed-length summary generation model, and obtain a dialogue summary target generation model according to the evaluation and optimization results;
[0089] The summary generation module is used to form a second dialogue summary template with the target dialogue corpus data and the second summary instruction, and use the dialogue summary target generation model to generate a dialogue summary of the target corpus data. The target dialogue corpus data is the dialogue data for which a dialogue summary is to be generated.
[0090] Unless otherwise specifically stated, the relative steps, numerical expressions, and numerical values of the components and steps set forth in these embodiments do not limit the scope of the present invention.
[0091] Each embodiment in this specification is described in a progressive manner. The key points of each embodiment are the differences from other embodiments. The same or similar parts among the embodiments can be referred to each other. For the system disclosed in the embodiments, since it corresponds to the method disclosed in the embodiments, the description is relatively simple, and the relevant parts can be referred to the description of the method part.
[0092] The units and method steps of each example described in combination with the embodiments disclosed in this article can be implemented by electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those of ordinary skill in the art can use different methods to implement the described functions for each specific application, but such implementation is not considered to exceed the scope of the present invention.
[0093] Those of ordinary skill in the art can understand that all or part of the steps in the above method can be completed by instructing relevant hardware through a program, and the program can be stored in a computer-readable storage medium, such as a read-only memory, a magnetic disk, or an optical disc, etc. Optionally, all or part of the steps of the above embodiments can also be implemented using one or more integrated circuits. Correspondingly, each module / unit in the above embodiments can be implemented in the form of hardware or in the form of a software functional module. The present invention is not limited to any specific form of the combination of hardware and software.
[0094] Finally, it should be noted that the above-described embodiments are only specific embodiments of the present invention, used to illustrate the technical solutions of the present invention, rather than to limit it. The protection scope of the present invention is not limited thereto. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: any person skilled in the art within the technical scope disclosed by the present invention can still modify the technical solutions recorded in the foregoing embodiments, or can easily think of changes, or make equivalent replacements for some of the technical features; and these modifications, changes, or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be covered by the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.
Claims
1. A method for automatically generating dialogue summaries based on fine-tuning a large language model, characterized in that: Include: Setting a first summary instruction and obtaining a training dialogue summary corpus dataset, wherein the first summary instruction is a prompt instruction label for instructing to generate a fixed-length dialogue summary, and the training dialogue summary corpus dataset includes a plurality of dialogue data samples and reference dialogue summaries, wherein each dialogue data sample is dialogue data information of consensus or decision-making communication carried out by at least one dialogue role around a specific dialogue task; The training dialogue summary corpus dataset is divided into a dialogue summary training dataset and a dialogue summary test dataset according to a preset ratio, the first summary instruction and the dialogue summary training dataset are combined to form a first dialogue summary template, and the first dialogue summary template is used to fine-tune the open source large language model to obtain a fixed-length summary generation model; Setting a second summary instruction, where the second summary instruction is a prompt instruction label for instructing to generate a conversation summary with no length limit, and combining the second summary instruction and the conversation summary test data set into a second conversation summary template, using the second conversation summary template to evaluate and optimize the fixed-length summary generation model, and obtaining a conversation summary target generation model according to the evaluation and optimization results; The target dialogue corpus data and the second summary instruction are combined into a second dialogue summary template, and a dialogue summary target generation model is used to generate a dialogue summary of the target corpus data, wherein the target dialogue corpus data is the dialogue data for which a dialogue summary is to be generated.
2. The method for automatically generating a dialogue summary based on a fine-tuned large language model according to claim 1, characterized in that: Get the conversation summary corpus dataset for training, including: Obtain a public conversation summary dataset, wherein the public conversation summary dataset includes: a conversation dataset constructed by translating daily monolingual conversations from a source language to a target language, and a conversation dataset constructed by collecting cross-language daily conversation pairs; Automatically generate task scenarios based on conversation summaries and select corresponding conversation datasets from public conversation summary datasets based on public conversation summary datasets, and use the selected conversation datasets as conversation summary corpus datasets for training.
3. The method for automatically generating a dialogue summary based on a fine-tuned large language model according to claim 1, characterized in that: Construct a dialogue summary template, including: The dialogue roles in the dialogue data are parsed, and the dialogue data is formatted according to the dialogue roles, and the dialogue summary is filled or inserted into the formatted dialogue data as a summary tag.
4. The method for automatically generating a dialogue summary based on a fine-tuned large language model according to claim 1, characterized in that: The open source large language model adopts a Chinese-English bilingual dialogue pre-training model or a lightweight pre-training language model.
5. The method for automatically generating a dialogue summary based on a fine-tuned large language model according to claim 1, characterized in that: The fixed-length summary generation model is evaluated and optimized using the second dialogue summary template, including: Evaluate the accuracy and conciseness of the model-generated summary based on the model-generated summary and the reference conversation summary; The first summary instruction and / or the second summary instruction are adjusted according to the evaluation result to determine the correctness of the set summary instruction.
6. The method for automatically generating a dialogue summary based on a fine-tuned large language model according to claim 5, characterized in that: Evaluate the accuracy of the model-generated summary, including: The reference conversation summary and the model-generated summary are split into sliding windows of n words to form n-grams of length n. The ROUGE-N index value used to evaluate the accuracy of the model-generated summary is calculated based on the number of n-grams that appear simultaneously in the reference conversation summary and the model-generated summary. Obtain the longest common subsequence between the reference conversation summary and the summary generated by the model, and calculate the ROUGE-L indicator value used to evaluate the accuracy of the summary generated by the model based on the longest common subsequence; The accuracy of the summary generated by the model is weightedly calculated based on the two indicator values of ROUGE-N and ROUGE-L. If the calculation result reaches the accuracy threshold, it is determined that the summary generated by the model meets the expected accuracy standard, and the dialogue summary target generation model is obtained.
7. The method for automatically generating a dialogue summary based on a fine-tuned large language model according to claim 6, characterized in that: Evaluate the conciseness of the model-generated summary, including: The length penalty factor is calculated according to the length of the reference dialogue summary and the length of the summary generated by the model. The BLEU index value used to evaluate the conciseness of the summary generated by the model is calculated based on the length penalty factor and the number of n-grams in the reference dialogue summary and the summary generated by the model. If the change in the BLEU index values of the reference dialogue summary and the summary generated by the model does not exceed the preset threshold, it is determined that the summary generated by the model meets the expected conciseness standard, and the dialogue summary target generation model is obtained.
8. A system for automatically generating dialogue summaries based on a fine-tuned large language model, characterized in that: It includes: instruction generation module, model fine-tuning module, model evaluation module and summary generation module, among which, an instruction generation module, configured to set a first summary instruction and obtain a training dialogue summary corpus data set, wherein the first summary instruction is a prompt instruction tag for instructing to generate a fixed-length dialogue summary, and the training dialogue summary corpus data set includes a plurality of dialogue data samples and reference dialogue summaries, wherein each dialogue data sample is dialogue data information of consensus or decision-making communication carried out by at least one dialogue role around a specific dialogue task; A model fine-tuning module is used to divide the conversation summary corpus dataset for training into a conversation summary training dataset and a conversation summary test dataset according to a preset ratio, form a first conversation summary template by pairing the first summary instruction with the conversation summary training dataset, and use the first conversation summary template to fine-tune the open source large language model to obtain a fixed-length summary generation model; A model evaluation module is used to set a second summary instruction, where the second summary instruction is a prompt instruction label indicating the generation of a dialogue summary with no length limit, and to form a second dialogue summary template with the second summary instruction and a dialogue summary test data set, and to evaluate and optimize the fixed-length summary generation model using the second dialogue summary template, and to obtain a dialogue summary target generation model according to the evaluation and optimization results; The summary generation module is used to combine the target dialogue corpus data and the second summary instruction into a second dialogue summary template, and use the dialogue summary target generation model to generate a dialogue summary of the target corpus data, wherein the target dialogue corpus data is the dialogue data for which the dialogue summary is to be generated.
9. An electronic device, characterized in that: include: at least one processor, and a memory coupled to the at least one processor; The memory stores a computer program, and the computer program can be executed by the at least one processor to implement the method according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed, the method according to any one of claims 1 to 7 can be implemented.