Dialogue abstract generation method and device and electronic equipment
By using scene type recognition and reinforcement learning to optimize the large language model, the problem of low efficiency in customer service dialogue summary generation for telecommunications operators has been solved, achieving efficient and accurate automated summary generation that can adapt to various customer service scenarios.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- CHINA MOBILE ONLINE SERVICES CO LTD
- Filing Date
- 2025-11-27
- Publication Date
- 2026-04-21
AI Technical Summary
In existing technologies, customer service dialogue summary generation solutions for telecommunications operators rely on manual operation, which suffers from low efficiency and high subjectivity. Furthermore, models based on deep neural networks struggle to achieve high-quality, on-demand output in complex and ever-changing customer service scenarios.
By acquiring dialogue text, a pre-trained scene type classification model is used to identify the target scene type. Combined with a large language model fine-tuned by reinforcement learning, a dialogue summary that conforms to the target summary template is generated.
It achieves automated generation of dialogue summaries, reduces manual intervention, improves generation efficiency and accuracy, ensures the objectivity and consistency of the summary content, and adapts to the format requirements of different scenarios.
Smart Images

Figure CN121901411A_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of natural language processing technology, specifically relating to a dialogue summary generation method, apparatus, and electronic device. Background Technology
[0002] In the customer service systems of telecommunications operators, users typically seek assistance through customer service hotlines when encountering issues such as network quality anomalies, disputes over package fees, or service interruptions. In related technologies, agents need to listen to the user's requests, extract key information, and manually input this information into a structured work order system according to a fixed template. This manual dialogue summary generation solution has several shortcomings. First, manual information extraction is easily influenced by subjective judgment, leading to information bias and affecting the accuracy and efficiency of subsequent processing. Second, facing the diverse application scenarios of intelligent customer service in operators, current solutions struggle to automatically match and output corresponding specified templates based on different scenarios, lacking the ability to adaptively output scenario-specific templates. Furthermore, although there are currently automatic generation models based on deep neural networks, most of these models rely on large amounts of labeled data for training, resulting in high training costs and long training times. At the same time, the learning capacity of deep neural networks is limited, making it difficult to fully capture key information in the dialogue, leading to insufficient model generalization ability and difficulty adapting to the complex and ever-changing needs of the customer service domain.
[0003] It is evident that current dialogue summarization schemes still face significant technical bottlenecks in the automated generation of dialogue summaries, particularly in the automatic identification of scene types and the adaptive output of corresponding templates. Therefore, providing a superior dialogue summarization scheme is essential. Summary of the Invention
[0004] This application provides a dialogue summarization method, apparatus, and electronic device that can solve the problems of low efficiency and insufficient accuracy in current dialogue summarization generation.
[0005] In a first aspect, embodiments of this application provide a dialogue summarization method, comprising: acquiring dialogue text between a user and an agent; inputting the dialogue text into a pre-trained scene type classification model, and outputting a target scene type corresponding to the dialogue text; determining a target summary template matching the target scene type based on the target scene type; inputting the dialogue text, the target summary template, and a first dialogue summary generation task prompt into a dialogue summary generation model, and guiding the dialogue summary generation model to output a first dialogue summary that corresponds to the dialogue text and conforms to the target summary template through the first dialogue summary generation task prompt; the dialogue summary generation model is obtained by optimizing a first large language model based on reinforcement learning fine-tuning.
[0006] Secondly, embodiments of this application provide a dialogue summarization generation apparatus, comprising: a first acquisition module for acquiring dialogue text between a user and an agent; a scene type classification processing module for inputting the dialogue text into a pre-trained scene type classification model and outputting a target scene type corresponding to the dialogue text; a determination module for determining a target summary template matching the target scene type based on the target scene type; and a dialogue summarization generation module for inputting the dialogue text, the target summary template, and a first dialogue summarization generation task prompt word into the dialogue summarization generation model, and guiding the dialogue summarization generation model to output a first dialogue summary that corresponds to the dialogue text and conforms to the target summary template through the first dialogue summarization generation task prompt word; the dialogue summarization generation model is obtained by optimizing a first large language model based on reinforcement learning fine-tuning.
[0007] Thirdly, embodiments of this application provide an electronic device including a processor; and a memory arranged to store computer-executable instructions configured to be executed by the processor to implement the steps of the dialogue summary generation method as described in the first aspect.
[0008] Fourthly, embodiments of this application provide a computer-readable storage medium for storing computer-executable instructions that, when executed by a processor, implement the steps of the dialogue summary generation method as described in the first aspect.
[0009] Fifthly, embodiments of this application provide a computer program product, the computer program product including a computer program that, when executed by a processor, implements the steps of the dialogue summary generation method as described in the first aspect.
[0010] In a sixth aspect, embodiments of this application provide a chip, the chip including a processor and a communication interface, the communication interface being coupled to the processor, the processor being configured to execute executable instructions to implement the steps of the dialogue summary generation method as described in the first aspect.
[0011] In this embodiment, the dialogue text between the user and the agent is acquired and input into a pre-trained scene type classification model. The model outputs the target scene type corresponding to the dialogue text. Based on the target scene type, a target summary template matching the target scene type is determined. Then, the dialogue text, the target summary template, and a first dialogue summary generation task prompt are input into the dialogue summary generation model. The first dialogue summary generation task prompt guides the model to output a first dialogue summary that corresponds to the dialogue text and conforms to the target summary template. The dialogue summary generation model is obtained by optimizing the first large language model based on reinforcement learning fine-tuning. It is evident that this technical solution achieves automatic scene type determination by accurately identifying the scene type of the dialogue text through the model. Based on this, a summary template is matched according to the determined scene type, and then the model generates a dialogue summary based on the matched summary template, achieving the effect of specifying a template according to the scene type and adaptively outputting the dialogue summary. Compared to the solution of manually summarizing dialogue summaries, this technical solution significantly reduces the manual intervention step, effectively avoids the problem of low efficiency in manual summarization, and improves the efficiency of dialogue summary generation. Meanwhile, this technical solution avoids subjective biases caused by differences in personal experience and understanding, ensuring that the generated dialogue summaries strictly adhere to unified standards and meet the format requirements of the corresponding scenarios. This ensures the objectivity and consistency of the dialogue summaries, further improving the standardization and scenario adaptability of the generated dialogue summaries. Furthermore, this technical solution employs reinforcement learning fine-tuning to optimize the large language model, thereby constructing a dialogue summary generation model. This model optimization strategy, even with limited labeled data, leverages data augmentation techniques to enable the model to fully learn from the data, effectively improving the overall performance of the model. Therefore, the optimized model possesses stronger capabilities in dialogue summaries. Thus, using this technical solution to generate dialogue summaries corresponding to dialogue text ensures the efficiency and accuracy of dialogue summaries generation, enabling agents to better understand and respond to users' actual needs, thereby improving the efficiency and accuracy of agent problem-solving. Attached Figure Description
[0012] Figure 1 This is a schematic block diagram of a dialogue summary generation system provided in an embodiment of this application; Figure 2 This is a flowchart illustrating a dialogue summary generation method provided in an embodiment of this application; Figure 3 This is a flowchart illustrating a dialogue summary generation method provided in another embodiment of this application; Figure 4 This is a scene classification error statistics chart provided in an embodiment of this application; Figure 5This is a visualization of a multi-dimensional evaluation result provided in an embodiment of this application; Figure 6 This is a schematic diagram of the structure of a dialogue summary generation device provided in an embodiment of this application; Figure 7 This is a schematic diagram of the structure of the electronic device provided in the embodiments of this application. Detailed Implementation
[0013] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0014] The terms "first," "second," etc., used in the specification and claims of this application are used to distinguish similar objects and not to describe a specific order or sequence. It should be understood that such use of data can be interchanged where appropriate so that embodiments of this application can be implemented in orders other than those illustrated or described herein, and the objects distinguished by "first," "second," etc., are generally of the same class and the number of objects is not limited; for example, a first object can be one or more. Furthermore, in the specification and claims, "and / or" indicates at least one of the connected objects, and the character " / " generally indicates that the preceding and following objects are in an "or" relationship.
[0015] It should be noted that the dialogue summary generation scheme provided in this application embodiment can be applied to various scenarios, such as customer service dialogue summary generation scenarios for telecommunications operators, work order generation scenarios for government hotlines, and work order generation scenarios for various service hotlines (such as financial services, application services, etc.). This application embodiment takes the customer service dialogue summary generation scenario of telecommunications operators as an example to describe in detail the dialogue summary generation method, device, and electronic equipment provided in this application embodiment.
[0016] In related technologies, the customer service dialogue summary generation schemes of telecommunications operators have the following problems: First, scenario selection mainly relies on manual methods. Specifically, agents first select scenarios from a scenario library based on the dialogue content with the user, thus generating a corresponding output template. Then, agents fill in the relevant elements in the determined output template according to the dialogue content. This method of relying on manual selection from a large number of scenarios is not only inefficient but also highly subjective and prone to errors. Therefore, the problem of automatically recommending scenarios based on dialogue content urgently needs to be solved.
[0017] Secondly, although current deep neural network-based models can automatically generate dialogue summaries, these generated contents are in open formats and cannot automatically generate corresponding content according to the requirements of specific scenario templates. Furthermore, due to the limited learning capacity of deep neural networks, even with a large amount of data labeled according to the template for each scenario, it is still difficult to achieve high-quality output as required. In addition, most current automatic dialogue summarization solutions based on large models adopt the SFT (Supervised Fine-Tuning) method. While this method can improve the model's adaptability and performance for specific tasks to some extent, it also has significant shortcomings. Specifically, real-world business scenarios are varied and complex, and fixed labeled data cannot fully cover all possible scenarios. This leads to inconsistent content quality when the model encounters unfamiliar or rarely occurring problem types, making it difficult to accurately capture key information. Moreover, the SFT method lacks the ability to understand deep semantic context, only extracting surface information, and often fails to effectively identify and correctly associate relevant information. Therefore, how to utilize limited labeled data to enable the model to fully learn key content and output according to the template format requirements is an urgent problem to be solved.
[0018] Therefore, this application provides a dialogue summary generation method. The following, in conjunction with the accompanying drawings, will describe in detail the dialogue summary generation method, apparatus, and electronic device provided by this application through specific embodiments and application scenarios.
[0019] Figure 1 This is a schematic block diagram of a dialogue summary generation system provided in an embodiment of this application, such as... Figure 1 As shown, the dialogue summary generation system may include a scene type recommendation module 110, a summary template adaptive generation module 120, a data intelligent analysis module 130, and a multi-agent-based training data construction module 140.
[0020] In one implementation, the dialogue summarization system may further include a speech recognition module 150. The speech recognition module may include an Automatic Speech Recognition (ASR) system. In a specific implementation, the dialogue between the user and the agent can be input into the speech recognition module in real time and processed into multi-turn dialogue text by the ASR. The multi-turn dialogue text is then used to input the scene type recommendation module 110 and the summary template adaptive generation module 120, respectively.
[0021] In one implementation, the scene type recommendation module 110 may include a large model and a scene type classification model. The large model may be an LLM (Large Language Model), and the scene type classification model may be a neural network-based classification model (e.g., the BERT model). In a specific implementation, the summarizing ability of the large model can be used to perform high-level semantic extraction on the input multi-turn dialogue text to obtain the core dialogue question. Then, the multi-turn dialogue text and the core dialogue question are input into the scene type classification model, and the output is the target scene type corresponding to the multi-turn dialogue text.
[0022] In one implementation, the dialogue summarization system may further include a summary template library 160. The summary template library can store the correspondence between scene types and summary templates. In a specific implementation, after obtaining the target scene type corresponding to the multi-turn dialogue text through the scene type recommendation module, the corresponding target summary template can be obtained by searching for the target scene type in the summary template library. The target summary template is used as input to the summary template adaptive generation module 120.
[0023] In one implementation, the summary template adaptive generation module 120 may include a preprocessing unit and a dialogue summary generation model. The dialogue summary generation model is obtained by optimizing the first large language model based on reinforcement learning fine-tuning. In a specific implementation, the preprocessing unit can sort the input multi-turn dialogue text according to the timeline to form ordered dialogue text. Based on the ordered dialogue text, the input target summary template, and preset prompts, a first dialogue summary generation task prompt is generated. Then, the dialogue text, the target summary template, and the first dialogue summary generation task prompt are input into the dialogue summary generation model. The first dialogue summary generation task prompt guides the dialogue summary generation model to output a first dialogue summary that corresponds to the dialogue text and conforms to the target summary template. The first dialogue summary is used to send to the agent.
[0024] After receiving the first conversation summary, the agent will output it directly if accepted. For example, the first conversation summary might be: "Upon verification, regarding the inquiry about 5G package change from number 13333333, we have contacted the customer, explained the package change process, and offered a two-month data allowance as a bonus. The customer has accepted."
[0025] If not adopted, the agent will manually edit the first dialogue summary to generate and output the second dialogue summary.
[0026] In one implementation, a data intelligence analysis module 130 can obtain a first dialogue summary and a second dialogue summary, respectively, and compare and analyze the first and second dialogue summaries to generate a multi-dimensional evaluation result of the dialogue summary generation model. The multi-dimensional evaluation result may include several of the following: incorrectly predicted entity names, unpredicted entity names, and incorrectly predicted scenario types. Optionally, after generating the multi-dimensional evaluation result of the dialogue summary generation model, it can be visualized.
[0027] In one implementation, a multi-dimensional evaluation result of the dialogue summarization generation model can be obtained through a multi-agent-based training data construction module 140, and then targeted training data for optimizing the dialogue summarization generation model can be constructed through agents based on the multi-dimensional evaluation result.
[0028] In this embodiment, by introducing a scene type recommendation module, the summarization capabilities of the large model and the classification capabilities of the small model can be leveraged to achieve automated scene type recommendation. By introducing an adaptive summary template generation module, dialogue summaries can be output according to the template requirements corresponding to the scene type. The model optimization strategy based on reinforcement learning fine-tuning can, with only a small amount of labeled data, utilize data augmentation techniques to enable the model to fully learn the knowledge in the data, effectively improving the overall performance of the model. Therefore, the optimized model has a stronger ability in dialogue summary generation. By constructing a data intelligent analysis module, the model's performance can be comprehensively evaluated according to predefined evaluation dimensions, revealing the model's shortcomings. Finally, by introducing a multi-agent training data construction module, combined with the analysis results of the data intelligent analysis module, model defects are identified and optimization directions are proposed. The system automatically selects and constructs the next training data, forming a Human-In-The-Loop (HITL, Human-Machine Collaboration) closed-loop optimization training system, enabling continuous model optimization.
[0029] Figure 2 This application illustrates a dialogue summary generation method according to an embodiment. This method can be executed by an electronic device, which may include a server and / or a terminal device, such as an in-vehicle terminal or a mobile terminal. In other words, the method can be executed by software or hardware installed on the electronic device, and the method includes the following steps: Step 202: Obtain the dialogue text between the user and the agent.
[0030] Step 204: Input the dialogue text into the pre-trained scene type classification model and output the target scene type corresponding to the dialogue text.
[0031] In practice, the summarizing capabilities of large models can be used to extract high-level semantics from the input dialogue text to obtain the core dialogue question. Then, the dialogue text and the core dialogue question are input into the scene type classification model, and the output is the target scene type corresponding to the dialogue text.
[0032] Step 206: Determine the target summary template that matches the target scene type based on the target scene type.
[0033] In the scenario of generating customer service dialogue summaries for telecommunications operators, the scenario types of dialogue text can include inquiry and consultation, customer praise, customer suggestions, customer complaints, installation and delivery, business opportunity reporting, touchpoint service quality, and broadband and internet TV fault reporting, etc.
[0034] Different scenario types correspond to different summary templates. As an example, the summary template for a query / consultation scenario is: "Upon verification, regarding the query / consultation for {product / service name} by {acceptance number}, we have / have not contacted the customer to explain the xx solution, and the customer accepts / does not accept it." The summary template for a customer suggestion scenario is: "Upon verification, regarding the xx situation of {suggestion type} {suggestion recipient} reflected by {acceptance number}, the solution is xx, we have / have not contacted the customer and explained it, and the customer accepts / does not accept it."
[0035] In practice, summary templates corresponding to each scenario type can be standardized and stored in a summary template library, indexed by scenario type. Therefore, after determining the target scenario type of the dialogue text, the corresponding target summary template can be obtained by searching for that target scenario type in the summary template library.
[0036] Step 208: Input the dialogue text, the target summary template, and the first dialogue summary generation task prompt into the dialogue summary generation model. Guide the dialogue summary generation model to output a first dialogue summary that corresponds to the dialogue text and conforms to the target summary template through the first dialogue summary generation task prompt.
[0037] Among them, the dialogue summarization generation model is obtained by optimizing the first language model based on reinforcement learning fine-tuning (ReFT).
[0038] In this embodiment, the dialogue text between the user and the agent is acquired and input into a pre-trained scene type classification model. The model outputs the target scene type corresponding to the dialogue text. Based on the target scene type, a target summary template matching the target scene type is determined. Then, the dialogue text, the target summary template, and a first dialogue summary generation task prompt are input into the dialogue summary generation model. The first dialogue summary generation task prompt guides the model to output a first dialogue summary that corresponds to the dialogue text and conforms to the target summary template. The dialogue summary generation model is obtained by optimizing the first large language model based on reinforcement learning fine-tuning. It is evident that this technical solution achieves automatic scene type determination by accurately identifying the scene type of the dialogue text through the model. Based on this, a summary template is matched according to the determined scene type, and then the model generates a dialogue summary based on the matched summary template, achieving the effect of specifying a template according to the scene type and adaptively outputting the dialogue summary. Compared to the solution of manually summarizing dialogue summaries, this technical solution significantly reduces the manual intervention step, effectively avoids the problem of low efficiency in manual summarization, and improves the efficiency of dialogue summary generation. Meanwhile, this technical solution avoids subjective biases caused by differences in personal experience and understanding, ensuring that the generated dialogue summaries strictly adhere to unified standards and meet the format requirements of the corresponding scenarios. This ensures the objectivity and consistency of the dialogue summaries, further improving the standardization and scenario adaptability of the generated dialogue summaries. Furthermore, this technical solution employs reinforcement learning fine-tuning to optimize the large language model, thereby constructing a dialogue summary generation model. This model optimization strategy, even with limited labeled data, leverages data augmentation techniques to enable the model to fully learn from the data, effectively improving the overall performance of the model. Therefore, the optimized model possesses stronger capabilities in dialogue summaries. Thus, using this technical solution to generate dialogue summaries corresponding to dialogue text ensures the efficiency and accuracy of dialogue summaries generation, enabling agents to better understand and respond to users' actual needs, thereby improving the efficiency and accuracy of agent problem-solving.
[0039] In one implementation, before obtaining the dialogue text between the user and the agent (i.e., step 202), a speech recognition system can be used to convert the dialogue voice between the user and the agent into multi-turn dialogue text.
[0040] As an example, a speech recognition system could be an Automatic Speech Recognition (ASR) system, which can convert the spoken conversation between the hotline user and the agent into a structured set of dialogue text, in the form of a multi-turn dialogue sequence. ,in, This represents a set of multi-turn dialogue text. Specifically, it consists of a series of rounds of user questions and agent responses. .
[0041] In this embodiment, the original dialogue speech is converted into a digital information stream suitable for subsequent text processing. This conversion not only enables the digital storage of speech data but also lays the foundation for subsequent automated processing. As an alternative, the speech recognition system can select different ASR models or services according to actual application needs to ensure conversion quality and real-time performance.
[0042] In one implementation, the training process of the scene type classification model may include the following steps A1-A6: Step A1: Obtain the second training dataset. The second training dataset includes multiple sets of second training data. Each set of second training data includes a second sample dialogue text and the second sample scene type corresponding to the second sample dialogue text.
[0043] In the context of customer service dialogue summary generation in telecommunications operators, there is currently a large amount of accumulated data, namely, data related to historical dialogue summary generation tasks. In this embodiment, a portion of the data can be cleaned from the accumulated database to construct a second training dataset. The cleaned data format is as follows: ,in , This represents the second sample of the dialogue text between the user and the agent. This indicates the second sample scenario type corresponding to the second sample dialogue text.
[0044] Step A2: For each set of second training data, use the large model to extract the core dialogue question corresponding to the second sample dialogue text.
[0045] In practical implementation, if the second sample dialogue text is used directly... Training a scene type classification model involves text that is too long and semantically complex, which increases the difficulty of model training. Therefore, this step can leverage the summarizing capabilities of a large model to perform high-level semantic extraction on the second sample dialogue text, thereby obtaining the core dialogue question. As an alternative, the extraction of core dialogue questions can be combined with different pre-trained large models, or a multi-model ensemble strategy can be adopted to enhance semantic understanding capabilities.
[0046] Step A3: Construct the third training data based on the second sample dialogue text, the second sample scene type, and the core dialogue question.
[0047] As an example, the third training data M .
[0048] Step A4: Input the third training data into the BERT model to be trained, and output the predicted scene type corresponding to the second sample dialogue text.
[0049] BERT (Bidirectional Encoder Representations from Transformers) is a pre-trained language model based on the Transformer architecture. During training, model performance can be optimized by adjusting the number of layers, attention mechanisms, or introducing regularization strategies. Alternatively, scene classification models can be replaced with other Transformer architectures or combined with lightweight models to adapt to different resource constraints.
[0050] Step A5: Determine the third loss result of the BERT model based on the second sample scene type, the predicted scene type, and the preset loss function.
[0051] Optionally, the preset loss function can be the cross-entropy loss function, used to quantify the comparison between the predicted probability distribution and the true label distribution; the greater the difference, the higher the loss value. As an example, the cross-entropy loss function... , This indicates the number of samples, specifically the number of third training data. Indicates the first The second sample scenario type of each sample Indicates the first The predicted scene type for each sample can be addressed by using weighted cross-entropy to deal with the problem of imbalanced scene categories.
[0052] Step A6: Based on the third loss result, update the model parameters of the BERT model to obtain the scene type classification model.
[0053] Specifically, the model parameters of the BERT model can be iteratively adjusted based on the third loss result until the iteration termination condition is met, at which point training stops, and the trained scene type classification model is obtained. Optionally, the iteration termination condition may include the convergence of the loss function and / or, the number of iterations reaching a preset number.
[0054] In this embodiment, a large model is used to perform high-level semantic extraction on long dialogue texts to obtain the core dialogue questions, which effectively reduces the complexity of the dialogue text and significantly highlights key information. Based on this, training data for scene type classification is generated using the extracted core dialogue questions, which helps improve the accuracy of the scene type classification model. Furthermore, by training a scene classifier based on the BERT model, automated recommendation of scene types for input dialogue texts can be achieved.
[0055] In one implementation, the training process of the dialogue summarization generation model may include the following steps B1-B3: Step B1: Obtain the first training dataset. The first training dataset includes multiple sets of first training data. Each set of first training data includes the first sample dialogue text, as well as the sample dialogue summary and sample entity information corresponding to the first sample dialogue text.
[0056] In one implementation, before executing step B1, a sample dataset can be obtained. This dataset includes the first sample dialogue text and a manually annotated sample dialogue summary corresponding to the first sample dialogue text. It is understood that the sample dialogue summary corresponding to the first sample dialogue text is obtained by determining the first sample scene type corresponding to the first sample dialogue text through a scene type classification model, thereby determining a sample summary template matching the first sample scene type, and then organizing human resources to annotate the first sample dialogue text according to the sample summary template.
[0057] As an example, the sample dataset , , This represents the first sample dialogue text between the user and the agent. This represents a manually annotated sample dialogue summary. In practice, the sample dataset can be split into two parts: one part is used for fine-tuning the training during the first-stage cold start of the model, enabling it to generate basically correct responses, which helps the second-stage reinforcement learning training converge quickly; the other part is used for the second-stage reinforcement learning training.
[0058] Considering that entity information plays a crucial role in dialogue summarization, its accuracy directly determines the precision and fidelity of the dialogue summary. To train the model's dialogue summarization capabilities, this embodiment employs a multi-task joint training approach to fine-tune the first large language model. Specifically, it enhances the model's capabilities through joint training of entity information extraction and dialogue summarization tasks. During this process, to accurately extract key information from the dialogue text between the user and the agent, the powerful generation capabilities of the large model can be leveraged to automatically construct a first subset for instruction fine-tuning training based on a subset of the sample dataset used for instruction fine-tuning training.
[0059] Specifically, based on a subset of the sample dataset used for fine-tuning training, prompt words can be used to guide the large model to output the sample entity information corresponding to the first sample dialogue text. As an example, the prompt words are as follows: "You are a customer service assistant for a telecommunications operator. Your task is to extract sample entity information from the sample dialogue summary, including the activity name, time, address, and mobile phone number, based on the first sample dialogue text and sample dialogue summary between the user and the agent of the telecommunications operator."
[0060] <dialogue_history> {dialogue} < / dialogue_history> <summarization> {summarization} < / summarization> <requirement> 1. Analyze the first sample dialogue text and the sample dialogue summary to extract key entity information. 2. Output in JSON format, for example, '{"Time": "Yesterday", "Address": "Beijing Chaoyang"}' < / requirement> output;”
[0061] Since the quality of training data directly determines the effectiveness of the trained model, choosing a good large model to construct the first subset for fine-tuning training is crucial. In practice, the DeepSeek-R1 series model can be selected as the model for constructing the first subset. The complete input of this model is: Through model processing, rules can be used to clean the generated data, resulting in the first subset. ,in, , This represents the first sample dialogue text between the user and the agent. This represents a manually annotated summary of sample dialogues. This represents sample entity information.
[0062] To further train the model's dialogue summarization capabilities, this embodiment employs a reinforcement learning algorithm to fine-tune the second language model. During this process, for the subset of the sample dataset used for reinforcement learning training, open-source tools such as Spacy can be used to extract sample entity information corresponding to the first sample dialogue text. Based on the subset of the sample dataset used for reinforcement learning training and the extracted sample entity information, a second subset for reinforcement learning training is then constructed.
[0063] Therefore, it can be understood that the first training dataset consists of a first subset and a second subset.
[0064] Step B2: Based on the first subset of the first training dataset, the first large language model is fine-tuned using a multi-task joint training method to obtain the second large language model.
[0065] Among them, the multi-task joint training method includes training the first language model through entity information extraction tasks and dialogue summary generation tasks.
[0066] Step B3: Based on the second subset within the first training dataset, fine-tune the second language model using a reinforcement learning algorithm to maximize the reward value of the target reward function, thereby obtaining the dialogue summary generation model.
[0067] In one implementation, the target reward function includes an edit distance reward function, an entity information reward function, and a template format reward function, with Kullback-Leibler Divergence introduced as a regularization term. The edit distance reward function measures the difference between the second predicted dialogue summary output by the second language model for the first sample dialogue text and the sample dialogue summary. The entity information reward function measures the accuracy of entity information extraction by the second language model. The template format reward function measures the accuracy of the template format of the second predicted dialogue summary.
[0068] This embodiment demonstrates a significant improvement in model capabilities even with limited training data. Specifically, by introducing entity information through data augmentation and combining it with multi-task joint training, the model can fully learn the knowledge within the data. Thus, by employing multi-task learning techniques, different tasks mutually reinforce each other, achieving multi-capability optimization of the model. Further fine-tuning through reinforcement learning enhances the model's ability to capture key information and improves output quality, thereby enhancing its generalization ability and robustness when facing complex and ever-changing customer service problems.
[0069] In one implementation, based on the first subset of the first training dataset, the first large language model is fine-tuned using a multi-task joint training method to obtain the second large language model (i.e., step B2), which can be executed as follows: steps B21-B24: Step B21: For each group of first training data in the first subset, construct first instruction fine-tuning samples and second instruction fine-tuning samples respectively based on the first training data.
[0070] The first instruction fine-tuning sample includes the first training data and entity information extraction task prompts. The second instruction fine-tuning sample includes the first sample dialogue text, the sample dialogue summary, the sample summary template corresponding to the first sample dialogue text, and the second dialogue summary generation task prompts. In specific implementation, the first sample dialogue text can be input into the scene type classification model, and the output will be the first sample scene type corresponding to the first sample dialogue text. Then, by searching for the first sample scene type in the summary template library, the corresponding sample summary template can be obtained.
[0071] Step B22 involves inputting the first instruction fine-tuning sample into the first large language model. The task prompt words for entity information extraction guide the first large language model to output the first predicted entity information corresponding to the first sample dialogue text. Thus, based on the sample entity information, the first predicted entity information, and the first loss function, the first loss result of the first large language model under the entity information extraction task is determined.
[0072] As an example, the prompts for entity information extraction tasks are as follows: “ input entity = """ Key entity information for generating a dialogue summary is extracted from the first sample dialogue text between the user and the agent. The first sample dialogue text { } <entity>: """ output entity = ".in, <entity>A special identifier for outputting the first predicted entity information.
[0073] In practice, the first loss function can be the cross-entropy loss function. As an example, the cross-entropy loss function... , This indicates the number of samples, specifically the number of dialogue texts in the first sample. This represents the model parameters for the entity information extraction task.
[0074] Step B23: Input the second instruction fine-tuning sample into the first large language model, and guide the first large language model to output the first predicted dialogue summary corresponding to the first sample dialogue text through the task prompt words generated by the second dialogue summary. Thus, based on the sample dialogue summary, the first predicted dialogue summary, and the second loss function, determine the second loss result of the first large language model under the dialogue summary generation task.
[0075] As an example, the task prompts for generating the second dialogue summary are as follows: " input summary = """ Based on the first sample dialogue text between the user and the agent, key entity information is extracted to summarize the dialogue. The first sample scene type is: {scene} The sample summary template is: {template} First sample dialogue text { } <summary>: """ output summary = ".in, <summary>A special identifier for outputting the first predicted dialogue summary.
[0076] In practice, the second loss function can be the cross-entropy loss function. As an example, the cross-entropy loss function... , This indicates the number of samples, specifically the number of dialogue texts in the first sample. This represents the model parameters for the second dialogue summary generation task.
[0077] Step B24: Based on the first and second loss results of the first language model, update the model parameters of the first language model to obtain the second language model.
[0078] In practical implementation, the loss function of multi-task joint training Therefore, by summing the first loss result and the second loss result, the loss result of multi-task joint training can be obtained.
[0079] Specifically, the model parameters of the first language model can be iteratively adjusted based on the loss results of multi-task joint training until the iteration termination condition is met, at which point training stops, and the second language model is obtained. Optionally, the iteration termination condition may include the convergence of the loss function and / or, the number of iterations reaching a preset number.
[0080] In this embodiment, a multi-task joint training strategy is adopted to collaboratively extract key entity information and generate dialogue summaries. By optimizing the model's comprehensive understanding ability through a joint loss function, the model can simultaneously ensure the accurate extraction of entity information and the effective summarization of dialogue summaries, thereby improving the completeness and accuracy of the generated dialogue summaries.
[0081] In one implementation, based on a second subset within the first training dataset, a reinforcement learning algorithm is used to fine-tune the second language model to maximize the reward result of the target reward function, resulting in a dialogue summarization generation model (i.e., step B3). This can be executed as follows: steps B31-B34. Step B31: For each set of first training data in the second subset, input the first training data, the sample summary template corresponding to the first sample dialogue text, and the first task prompt word into the second large language model. Guide the second large language model to output the second predicted entity information and the second predicted dialogue summary corresponding to the first sample dialogue text through the first task prompt word.
[0082] Understandably, the second language model is obtained by fine-tuning the first language model using a multi-task joint training method. Therefore, the second language model can process entity information extraction and dialogue summary generation tasks based on the corresponding prompt words (first task prompt words), thereby obtaining the second predicted entity information and the second predicted dialogue summary corresponding to the first sample dialogue text.
[0083] Step B32: Determine the entity reward result of the second language model based on the sample entity information, the second predicted entity information, and the preset entity reward function; determine the edit distance reward result of the second language model based on the sample dialogue summary, the second predicted dialogue summary, and the preset edit distance reward function; and determine the template format reward result of the second language model based on the sample summary template, the second predicted dialogue summary, and the preset template format reward function.
[0084] In practice, the edit distance reward function can be: .in, Represents a sample dialogue summary Summary of the Second Prediction Dialogue The smaller the edit distance between the second predicted dialogue summary generated by the second largest language model and the standard output (i.e., the sample dialogue summary), the larger the reward is given, i.e., the edit distance reward result. The larger the edit distance, the less reward is given; conversely, the greater the edit distance, the less reward is given.
[0085] The entity reward function can be: .in, The number of entities representing the sample entity information (i.e., the standard output). This indicates the number of entities in the second predicted entity information (i.e., the model prediction output). This indicates the number of correctly predicted entities in the second predicted entity information. This represents the number of incorrectly predicted entities in the second predicted entity information. According to the entity reward function, the more correctly predicted entities the model outputs, the greater the reward, i.e., the entity reward result. The larger the value, the lower the reward; conversely, the fewer entities the model correctly predicts in the output, the less reward is given. Therefore, this entity information reward function can measure the accuracy of entity information extraction by the second-largest language model.
[0086] The template-formatted reward function can be: By introducing a template-based reward function, the second-largest language model can be prompted to accurately output dialogue summaries according to the summary templates for the corresponding scene types.
[0087] Step B33: Based on the entity reward result, edit distance reward result, template format reward result, and KL divergence of the second largest language model, determine the reward result of the objective reward function of the second largest language model.
[0088] In practical implementation, the target reward function can be: ,in, , For hyperparameters, The coefficient of the KL divergence.
[0089] Step B34: Optimize the PPO algorithm according to the near-end strategy to update the model parameters of the second language model until the reward result of the target reward function reaches the set condition or the number of updates reaches the set number, thus obtaining the dialogue summary generation model.
[0090] Among them, PPO (Proximal Policy Optimization) is a reinforcement learning algorithm based on policy gradients. It updates the policy through proximal policy optimization to achieve stable and efficient training results. In algorithms such as PPO, KL divergence is used to control the degree of deviation between the old and new policies. If the KL divergence is too large (overly aggressive policy update), the parameter update step size will be reduced; if it is too small (insufficient policy change), a larger step size is allowed.
[0091] During the execution of the PPO algorithm, the generalized advantage estimation (GAE) scheme is used to calculate the reinforcement learning advantage: ,in, As a discount factor for rewards, This is the discount factor for temporal differences, which are defined as follows: Value networks are established through policy models. (That is, the second largest language model that needs to be updated) After the output of the last hidden state layer, a linear feedforward network (FFN) is added to construct it. , It is a value network for state The predicted value. Regarding the estimate of returns, It is the target value of actual return. That is, the estimate of the return can be the sum of the generalized advantage estimate and the value estimate. The policy loss function of the PPO algorithm is as follows: .
[0092] The value loss function is as follows: .
[0093] The final loss function is: , These are the coefficients of the value loss function.
[0094] The process of fine-tuning the second language model using reinforcement learning algorithms involves multiple loops. In each loop, the second language model generates a response (including second predicted entity information and a second predicted dialogue summary corresponding to the first sample dialogue text), receives evaluation from the target reward function, and adjusts its behavior accordingly. Based on the reward received, the model updates its behavioral policy to improve future outputs. Specifically, if the second language model achieves a high score, it reinforces the current policy; if it achieves a low score, it adjusts the policy and tries new methods. This loop is repeated continuously, thereby continuously improving the performance of the second language model.
[0095] In this embodiment, the second language model is fine-tuned and optimized based on the PPO algorithm. A weighted comprehensive reward function combining entity reward, edit distance reward, and template format reward is designed, and KL divergence is introduced as a regularization term, which helps to improve the format standardization and information accuracy of the dialogue summary generation model output.
[0096] In one implementation, the agent can review the first dialogue summary output by the dialogue summary generation model and then decide whether to accept or reject it. If the agent accepts it, it is output directly. If not, the agent manually edits the first dialogue summary to generate a second dialogue summary and outputs it.
[0097] Therefore, as Figure 3 As shown, the output of the dialogue summarization model can be evaluated in multiple dimensions by performing the following steps 302-306, and targeted training data can be constructed based on the multi-dimensional evaluation results to continuously optimize the scene type classification model and the dialogue summarization model.
[0098] Step 302: Obtain the second dialogue summary. The second dialogue summary is generated by the agent after editing the first dialogue summary.
[0099] Step 304: Compare and analyze the first dialogue summary and the second dialogue summary to generate a multi-dimensional evaluation result of the dialogue summary generation model.
[0100] The multi-dimensional evaluation results can include several of the following: the entity names that were predicted incorrectly, the entity names that were not predicted, and the scenario types that were predicted incorrectly.
[0101] In practical implementation, the entity reward function, edit distance reward function, and template format reward function defined during the fine-tuning of the second large language model using the aforementioned reinforcement learning algorithm can be used to compare and analyze the first and second dialogue summaries. After comparing and analyzing each pair of first and second dialogue summaries, the total reward is as follows: .
[0102] When the total reward is less than a set threshold, the first dialogue summary can be identified as complex sample data. By collecting and accumulating complex sample data, the dialogue summary generation model can be continuously optimized.
[0103] After accumulating complex sample data according to the above scheme, the information corresponding to each data point can include: the summary content of the model prediction output, the summary content after agent editing, and the edit distance reward. Template format rewards Physical rewards Entity names that are predicted incorrectly Unpredicted entity names Predicting scenario types Scene types after seat editing .
[0104] In this context, incorrectly predicted entity names refer to entity names that appeared in the first dialogue summary output by the model but did not appear in the second dialogue summary edited by the agent, meaning the agent deleted the corresponding entity name. Unpredicted entity names refer to entity names that did not appear in the first dialogue summary but appeared in the second dialogue summary.
[0105] As an example, the multi-dimensional evaluation results could include the set of entity names from the first dialogue summary. The set of entity names in the second dialogue summary The set of correctly predicted entity names (including entity names that appear in both of the above sets). The set of entity names that were predicted incorrectly and the set of unpredicted entity names This provides a comprehensive evaluation of the quality of the model's predictive output and offers quantitative indicators across multiple dimensions.
[0106] To provide optimization direction for the scene type classification model, we can statistically analyze the error rate of different scene types from the complex sample data accumulated above, that is, the scene type predicted by the model ≠ the scene type after the agent edits. , Indicates the type of scene predicted by the model. This indicates the scene type after the agent's seat has been edited. As an example, Figure 4 The chart shows the statistics of scene classification errors over a certain period of time.
[0107] Optionally, after obtaining the multi-dimensional evaluation results of the dialogue summarization generation model, they can be visualized separately. For example... Figure 5 As shown, (a) presents error statistics for different scenarios, including entity errors, field errors, and template errors. (b) presents statistics on the frequency of different entity errors. (c) presents statistics on unpredicted high-frequency entities.
[0108] Step 306: Based on the multi-dimensional evaluation results, construct targeted training data for optimizing the dialogue summary generation model through the agent.
[0109] In practical implementation, based on the multi-dimensional evaluation results, a large language model can be used to generate an analysis report. According to the statistical content in step 304, the analysis report can include the following four parts: statistics on scene classification errors, optimization directions, and data solutions required for optimization; statistics on content error types under different scenes, optimization directions, and data solutions required for optimization; the model's prediction of high-frequency erroneous entities, optimization directions, and data solutions required for optimization; and the model's failure to predict the distribution of high-frequency entities, optimization directions, and data solutions required for optimization. The prompts used to guide the large language model in generating the analysis report can be: Role: Optimization Analyst for Telecom Operator Dialogue Summary Generation System Background: Based on the data accumulated by the dialogue summary generation system, four core statistical components have been completed through the data intelligence analysis module: 1. Distribution of scene classification errors 2. Error type distribution in different scenarios (format / entity / logic errors) 3. List of frequently mispredicted entities by the model 4. The model failed to predict the distribution of high-frequency entities. Database fields include: - model_prediction (model-predicted dialogue summary content) - edited_content (Summarized dialogue content edited by the agent) - edit_distance_reward (Edit distance reward) - template_format_reward (template format reward) - entity_reward (entity reward) - predicted_error_entities (names of predicted error entities) - missed_entities (Entity names not predicted) - predicted_scene (predicted scene type) - edited_scene (Scene type after agent editing) Task: Generate a four-part structured analysis report, each containing three core modules: 1. Visual description of statistical results (based on existing analysis results) 2. Targeted optimization directions, providing specific data selection schemes. 3. Optimize the required data collection plan. Output format requirements: ### [Title] **Statistical Overview** [Visualized text description] **Optimization Directions** **Optimize data acquisition scheme**.
[0110] For the generated analysis report, the Text-to-SQL large language model can be used to generate the corresponding SQL statement, thereby querying the corresponding data from the accumulated data and processing it into training data format to participate in the next round of model optimization training.
[0111] In practical implementation, based on the accumulated data, targeted training data for optimizing the dialogue summarization generation model can be constructed using intelligent agents. These agents can include agents for constructing data based on scene classification errors, content errors, high-frequency entity prediction errors, and high-frequency unpredicted entity errors. This multi-agent automated data construction process effectively improves the relevance and quality of training data, facilitating continuous iteration and optimization of the model in real-world business applications.
[0112] Training data needed for scene classification error optimization can be extracted from the database. An agent is constructed using scene classification error data constructed from prompt words. This agent generates SQL statements, which are then executed using the available tool `Execute_sql` to select data. Finally, the provided tool `data_process0` processes the selected data into the format required for training. As an example, the prompt words are as follows: Role: Optimization Analyst for Telecom Operator Dialogue Summary Generation System Background: Based on the data accumulated by the dialogue summarization generation system, the optimization direction for scene classification errors, and the data collection scheme required for optimization, this study aims to select training data from the accumulated database that can be used to optimize the model. Optimization direction for scene classification errors: {secene_error} Optimize the required data acquisition scheme: {secene_polocy} Database fields include: - model_prediction (model-predicted dialogue summary content) - edited_content (Summarized dialogue content edited by the agent) - edit_distance_reward (Edit distance reward) - template_format_reward (template format reward) - entity_reward (entity reward) - predicted_error_entities (names of predicted error entities) - missed_entities (Entity names not predicted) - predicted_scene (predicted scene type) - edited_scene (Scene type after agent editing) Tools: Execute_sql: A data selection tool that executes SQL statements and selects data that meets specific requirements. data_process0: A data post-processing tool that selects data and processes it into text of a specified format. Task: It outputs executable Text-to-SQL statements that select data from the database that meets the requirements and process it into the corresponding format. You can follow the steps as follows: Think -- Act -- Observation.
[0113] Training data for content error optimization can be extracted from the database. An agent is constructed based on the content error data using prompt words. This agent generates SQL statements, which are then executed using the available tool `Execute_sql` to select data. Finally, the provided tool `data_process0` processes the selected data into the format required for training. As an example, the prompt words are as follows: Role: Optimization Analyst for Telecom Operator Dialogue Summary Generation System Background: Based on the data accumulated by the dialogue summarization system, this study optimizes the model by selecting training data from the accumulated database according to three error types (formatting errors, entity errors, and field content errors) and the required data acquisition scheme. Error type optimization direction: {type_error} Optimize the required data acquisition scheme: {type_polocy} Database fields include: - model_prediction (model-predicted dialogue summary content) - edited_content (Summarized dialogue content edited by the agent) - edit_distance_reward (Edit distance reward) - template_format_reward (template format reward) - entity_reward (entity reward) - predicted_error_entities (names of predicted error entities) - missed_entities (Entity names not predicted) - predicted_scene (predicted scene type) - edited_scene (Scene type after agent editing) Tools: Execute_sql: A data selection tool that executes SQL statements and selects data that meets specific requirements. data_process0: A data post-processing tool that selects data and processes it into text of a specified format. Task: It outputs executable Text-to-SQL statements that select data from the database that meets the requirements and process it into the corresponding format. You can follow the steps as follows: Think -- Act -- Observation.
[0114] Training data for optimizing high-frequency entity prediction errors can be extracted from the database. An agent is constructed using this high-frequency entity prediction error data, which generates SQL statements. The available tool `Execute_sql` is used to execute these SQL statements to select data. Finally, the provided tool `data_process0` processes the selected data into the format required for training. As an example, the prompts are as follows: Role: Optimization Analyst for Telecom Operator Dialogue Summary Generation System Background: Based on the data accumulated by the dialogue summarization system, the optimization direction for high-frequency entity prediction errors, and the data acquisition scheme required for optimization, this study aims to select training data from the accumulated database that can be used to optimize the model. Optimization direction for high-frequency entity prediction errors: {ent_error} Optimize the required data acquisition scheme: {ent_polocy} Database fields include: - model_prediction (model-predicted dialogue summary content) - edited_content (Summarized dialogue content edited by the agent) - edit_distance_reward (Edit distance reward) - template_format_reward (template format reward) - entity_reward (entity reward) - predicted_error_entities (names of predicted error entities) - missed_entities (Entity names not predicted) - predicted_scene (predicted scene type) - edited_scene (Scene type after agent editing) Tools: Execute_sql: A data selection tool that executes SQL statements and selects data that meets specific requirements. data_process0: A data post-processing tool that selects data and processes it into text of a specified format. Task: It outputs executable Text-to-SQL statements that select data from the database that meets the requirements and process it into the corresponding format. You can follow the steps as follows: Think -- Act -- Observation.
[0115] Training data for optimizing high-frequency unpredictable entity errors can be extracted from the database. An agent is constructed based on this high-frequency unpredictable entity error data using prompt words. This agent generates SQL statements, which are then executed using the available tool `Execute_sql` to select data. Finally, the provided tool `data_process0` processes the selected data into the format required for training. As an example, the prompt words are as follows: Role: Optimization Analyst for Telecom Operator Dialogue Summary Generation System Background: Based on the data accumulated by the dialogue summarization system, the optimization direction for high-frequency entity prediction errors, and the data acquisition scheme required for optimization, this study aims to select training data from the accumulated database that can be used to optimize the model. Optimization direction for high-frequency entity prediction errors: {ent_error} Optimize the required data acquisition scheme: {ent_polocy} Database fields include: - model_prediction (model-predicted dialogue summary content) - edited_content (Summarized dialogue content edited by the agent) - edit_distance_reward (Edit distance reward) - template_format_reward (template format reward) - entity_reward (entity reward) - predicted_error_entities (names of predicted error entities) - missed_entities (Entity names not predicted) - predicted_scene (predicted scene type) - edited_scene (Scene type after agent editing) Tools: Execute_sql: A data selection tool that executes SQL statements and selects data that meets specific requirements. data_process0: A data post-processing tool that selects data and processes it into text of a specified format. Task: It outputs executable Text-to-SQL statements that select data from the database that meets the requirements and process it into the corresponding format. You can follow the steps as follows: Think -- Act -- Observation.
[0116] In this embodiment, by introducing editorial feedback from agents on the dialogue summaries generated by the model, and by calculating multi-dimensional evaluation results, complex sample data can be identified and filtered to promote targeted learning of the model for challenging samples. Intelligent agents automatically generate intelligent analysis reports, constructing diverse training data including scene classification errors, content errors, high-frequency entity prediction errors, and high-frequency unpredicted entity errors. Each agent automatically generates and executes SQL statements to complete the filtering and processing of training data, forming a closed-loop system for continuous iterative optimization of the model, ensuring the model's adaptability and stability under diverse business needs across multiple scenarios.
[0117] In summary, the embodiments of this application achieve efficient conversion of multi-turn dialogue text by fully utilizing speech recognition technology, and combine it with a large language model for high-level semantic extraction and core question refinement, significantly reducing text complexity and improving the accuracy of scene classification. Simultaneously, the BERT-based scene classification model automatically recommends scene types, reducing the subjectivity and workload of manual operations. By integrating entity information extraction and dialogue summary generation through a multi-task joint training strategy, the model's ability to understand and generate key information is enhanced. PPO reinforcement learning fine-tuning is introduced, combining edit distance, entity accuracy, and template format rewards to optimize the model's output quality and format standardization. Through edit feedback collection and complex sample accumulation mechanisms, continuous monitoring and improvement of model performance are achieved. Through multi-agent data construction, optimization reports and targeted training data are automatically generated, constructing an automated closed-loop optimization system to ensure the model continuously evolves in practical applications.
[0118] These technological features work synergistically to significantly improve the efficiency and accuracy of dialogue summary generation in the dialogue summary generation system, optimizing user experience and customer service workflows. Specifically, the system can quickly understand and categorize customer complaints, accurately capture user needs, and automatically output dialogue summary content that conforms to scenario templates, thereby reducing customer wait time and improving customer satisfaction and brand loyalty. Simultaneously, automated work order generation and allocation improve agent efficiency, allowing them to focus on handling complex issues. Intelligent data analysis supports internal management optimization, promotes product and service improvements, and enhances enterprise competitiveness. This technical solution effectively overcomes the shortcomings of traditional manual modes and existing deep network models in dialogue summary extraction, achieving higher generalization ability and adaptability, and meeting the diverse and complex practical needs of telecommunications operators in the field of intelligent customer service.
[0119] It should be noted that the dialogue summarization method provided in this application can be executed by a dialogue summarization generation device or a control module within that device for executing the dialogue summarization method. This application uses the example of a dialogue summarization device executing the dialogue summarization method to illustrate the dialogue summarization device provided in this application.
[0120] Figure 6 This is a schematic diagram of a dialogue summary generation device provided in an embodiment of this application. Figure 6 As shown, the dialogue summary generation device includes: a first acquisition module 610, a scene type classification processing module 620, a determination module 630, and a dialogue summary generation module 640.
[0121] The first acquisition module 610 is used to acquire the dialogue text between the user and the agent; the scene type classification processing module 620 is used to input the dialogue text into a pre-trained scene type classification model and output the target scene type corresponding to the dialogue text; the determination module 630 is used to determine the target summary template that matches the target scene type; the dialogue summary generation module 640 is used to input the dialogue text, the target summary template, and the first dialogue summary generation task prompt words into the dialogue summary generation model, and guide the dialogue summary generation model to output a first dialogue summary that corresponds to the dialogue text and conforms to the target summary template through the first dialogue summary generation task prompt words; the dialogue summary generation model is obtained by optimizing the first large language model based on reinforcement learning fine-tuning.
[0122] In one implementation, the training process of the dialogue summarization generation model includes: acquiring a first training dataset; the first training dataset includes multiple sets of first training data, each set of first training data including a first sample dialogue text, and sample dialogue summaries and sample entity information corresponding to the first sample dialogue text; based on a first subset within the first training dataset, fine-tuning the first large language model using a multi-task joint training method to obtain a second large language model; the multi-task joint training method includes training the first large language model through entity information extraction tasks and dialogue summarization generation tasks; based on a second subset within the first training dataset, fine-tuning the second large language model using a reinforcement learning algorithm to maximize the reward value of the target reward function, thereby obtaining the dialogue summarization generation model.
[0123] In one implementation, a second language model is obtained by fine-tuning a first language model using a multi-task joint training method based on a first subset of the first training dataset. This includes: for each set of first training data in the first subset, constructing first and second instruction fine-tuning samples based on the first training data; the first instruction fine-tuning sample includes the first training data and entity information extraction task prompts; the second instruction fine-tuning sample includes first sample dialogue text, a sample dialogue summary, a sample summary template corresponding to the first sample dialogue text, and second dialogue summary generation task prompts; the first instruction fine-tuning sample is input into the first language model, and the first language model is guided by the entity information extraction task prompts. The system outputs the first predicted entity information corresponding to the first sample dialogue text; based on the sample entity information, the first predicted entity information, and the first loss function, it determines the first loss result of the first language model under the entity information extraction task; it inputs the second instruction fine-tuning sample into the first language model, and guides the first language model to output the first predicted dialogue summary corresponding to the first sample dialogue text through the second dialogue summary generation task prompt words; based on the sample dialogue summary, the first predicted dialogue summary, and the second loss function, it determines the second loss result of the first language model under the dialogue summary generation task; based on the first loss result and the second loss result of the first language model, it updates the model parameters of the first language model to obtain the second language model.
[0124] In one implementation, based on a second subset within the first training dataset, a reinforcement learning algorithm is used to fine-tune a second large language model to maximize the reward result of the target reward function, resulting in a dialogue summarization generation model. This includes: for each set of first training data in the second subset, inputting the first training data, a sample summary template corresponding to the first sample dialogue text, and a first task prompt word into the second large language model; guiding the second large language model to output second predicted entity information and a second predicted dialogue summary corresponding to the first sample dialogue text through the first task prompt word; determining the entity reward result of the second large language model based on the sample entity information, the second predicted entity information, and a preset entity reward function; and... Based on the sample dialogue summary, the second predicted dialogue summary, and the preset edit distance reward function, the edit distance reward result of the second largest language model is determined; and based on the sample summary template, the second predicted dialogue summary, and the preset template format reward function, the template format reward result of the second largest language model is determined; based on the entity reward result, the edit distance reward result, the template format reward result, and the KL divergence of the second largest language model, the reward result of the target reward function of the second largest language model is determined; the model parameters of the second largest language model are updated according to the proximal policy optimization PPO algorithm until the reward result of the target reward function reaches the set condition or the number of updates reaches the set number, thus obtaining the dialogue summary generation model.
[0125] In one implementation, the dialogue summary generation device further includes: a second acquisition module, an evaluation result generation module, and a training data construction module.
[0126] The second acquisition module is used to acquire a second dialogue summary; the second dialogue summary is generated by the agent after editing the first dialogue summary; the evaluation result generation module is used to compare and analyze the first and second dialogue summaries to generate a multi-dimensional evaluation result of the dialogue summary generation model; the multi-dimensional evaluation result includes several of the following: incorrectly predicted entity names, unpredicted entity names, and incorrectly predicted scenario types; the training data construction module is used to construct targeted training data for optimizing the dialogue summary generation model through the agent based on the multi-dimensional evaluation results.
[0127] In one implementation, the training process of the scene type classification model includes: acquiring a second training dataset; the second training dataset includes multiple sets of second training data, each set including a second sample dialogue text and a second sample scene type corresponding to the second sample dialogue text; for each set of second training data, using a large model to extract the core dialogue question corresponding to the second sample dialogue text; constructing third training data based on the second sample dialogue text, the second sample scene type, and the core dialogue question; inputting the third training data into the BERT model to be trained, outputting the predicted scene type corresponding to the second sample dialogue text; determining the third loss result of the BERT model based on the second sample scene type, the predicted scene type, and a preset loss function; updating the model parameters of the BERT model based on the third loss result to obtain the scene type classification model.
[0128] In this embodiment, the dialogue text between the user and the agent is acquired and input into a pre-trained scene type classification model. The model outputs the target scene type corresponding to the dialogue text. Based on the target scene type, a target summary template matching the target scene type is determined. Then, the dialogue text, the target summary template, and a first dialogue summary generation task prompt are input into the dialogue summary generation model. The first dialogue summary generation task prompt guides the model to output a first dialogue summary that corresponds to the dialogue text and conforms to the target summary template. The dialogue summary generation model is obtained by optimizing the first large language model based on reinforcement learning fine-tuning. It is evident that this technical solution achieves automatic scene type determination by accurately identifying the scene type of the dialogue text through the model. Based on this, a summary template is matched according to the determined scene type, and then the model generates a dialogue summary based on the matched summary template, achieving the effect of specifying a template according to the scene type and adaptively outputting the dialogue summary. Compared to the solution of manually summarizing dialogue summaries, this technical solution significantly reduces the manual intervention step, effectively avoids the problem of low efficiency in manual summarization, and improves the efficiency of dialogue summary generation. Meanwhile, this technical solution avoids subjective biases caused by differences in personal experience and understanding, ensuring that the generated dialogue summaries strictly adhere to unified standards and meet the format requirements of the corresponding scenarios. This ensures the objectivity and consistency of the dialogue summaries, further improving the standardization and scenario adaptability of the generated dialogue summaries. Furthermore, this technical solution employs reinforcement learning fine-tuning to optimize the large language model, thereby constructing a dialogue summary generation model. This model optimization strategy, even with limited labeled data, leverages data augmentation techniques to enable the model to fully learn from the data, effectively improving the overall performance of the model. Therefore, the optimized model possesses stronger capabilities in dialogue summaries. Thus, using this technical solution to generate dialogue summaries corresponding to dialogue text ensures the efficiency and accuracy of dialogue summaries generation, enabling agents to better understand and respond to users' actual needs, thereby improving the efficiency and accuracy of agent problem-solving.
[0129] The dialogue summary generation device in this application embodiment can be a device, or a component, integrated circuit, or chip in a terminal. The device can be a mobile electronic device or a non-mobile electronic device. For example, mobile electronic devices can be mobile phones, tablets, laptops, PDAs, in-vehicle electronic devices, wearable devices, ultra-mobile personal computers (UMPCs), netbooks, or personal digital assistants (PDAs), etc., while non-mobile electronic devices can be servers, network-attached storage (NAS), personal computers (PCs), televisions (TVs), ATMs, or self-service machines, etc. This application embodiment does not impose specific limitations.
[0130] The dialogue summary generation device in this application embodiment can be a device with an operating system. This operating system can be Android, iOS, or other possible operating systems; this application embodiment does not specifically limit it.
[0131] The dialogue summation device provided in this application embodiment can achieve... Figures 2 to 5 The various processes implemented in the method embodiments are not described in detail here to avoid repetition.
[0132] Based on the same technical concept, embodiments of this application also provide an electronic device for performing the above-described dialogue summary generation method. Figure 7 This is a schematic diagram of the structure of an electronic device to implement various embodiments of this application. The electronic device can vary significantly due to differences in configuration or performance, and may include a processor 710, a communications interface 720, a memory 730, and a communication bus 740. The processor 710, communications interface 720, and memory 730 communicate with each other via the communication bus 740. The processor 710 can call a computer program stored in the memory 730 and executable on the processor 710 to perform the following steps: The process involves: acquiring the dialogue text between the user and the agent; inputting the dialogue text into a pre-trained scene type classification model, which outputs the target scene type corresponding to the dialogue text; determining a target summary template that matches the target scene type based on the target scene type; inputting the dialogue text, the target summary template, and the first dialogue summary generation task prompt into the dialogue summary generation model, which is then guided by the first dialogue summary generation task prompt to output a first dialogue summary that corresponds to the dialogue text and conforms to the target summary template; and optimizing the first language model based on reinforcement learning fine-tuning.
[0133] In this embodiment, the dialogue text between the user and the agent is acquired and input into a pre-trained scene type classification model. The model outputs the target scene type corresponding to the dialogue text. Based on the target scene type, a target summary template matching the target scene type is determined. Then, the dialogue text, the target summary template, and a first dialogue summary generation task prompt are input into the dialogue summary generation model. The first dialogue summary generation task prompt guides the model to output a first dialogue summary that corresponds to the dialogue text and conforms to the target summary template. The dialogue summary generation model is obtained by optimizing the first large language model based on reinforcement learning fine-tuning. It is evident that this technical solution achieves automatic scene type determination by accurately identifying the scene type of the dialogue text through the model. Based on this, a summary template is matched according to the determined scene type, and then the model generates a dialogue summary based on the matched summary template, achieving the effect of specifying a template according to the scene type and adaptively outputting the dialogue summary. Compared to the solution of manually summarizing dialogue summaries, this technical solution significantly reduces the manual intervention step, effectively avoids the problem of low efficiency in manual summarization, and improves the efficiency of dialogue summary generation. Meanwhile, this technical solution avoids subjective biases caused by differences in personal experience and understanding, ensuring that the generated dialogue summaries strictly adhere to unified standards and meet the format requirements of the corresponding scenarios. This ensures the objectivity and consistency of the dialogue summaries, further improving the standardization and scenario adaptability of the generated dialogue summaries. Furthermore, this technical solution employs reinforcement learning fine-tuning to optimize the large language model, thereby constructing a dialogue summary generation model. This model optimization strategy, even with limited labeled data, leverages data augmentation techniques to enable the model to fully learn from the data, effectively improving the overall performance of the model. Therefore, the optimized model possesses stronger capabilities in dialogue summaries. Thus, using this technical solution to generate dialogue summaries corresponding to dialogue text ensures the efficiency and accuracy of dialogue summaries generation, enabling agents to better understand and respond to users' actual needs, thereby improving the efficiency and accuracy of agent problem-solving.
[0134] The specific execution steps can be found in the various steps of the above-described dialogue summary generation method embodiment, and can achieve the same technical effect. To avoid repetition, they will not be repeated here.
[0135] It should be noted that the electronic devices in the embodiments of this application include: servers, terminals, or other devices besides terminals.
[0136] The above electronic device structure does not constitute a limitation on the electronic device. An electronic device may include more or fewer components than illustrated, or combine certain components, or arrange them differently. For example, an input unit may include a Graphics Processing Unit (GPU) and a microphone, and a display unit may use a liquid crystal display (LCD), organic light-emitting diode (OLED), or other similar display panels. User input units include at least one of a touch panel and other input devices. A touch panel is also called a touchscreen. Other input devices may include, but are not limited to, physical keyboards, function keys (such as volume control buttons, power buttons, etc.), trackballs, mice, and joysticks, which will not be elaborated further here.
[0137] Memory can be used to store software programs and various data. Memory can primarily include a first storage area for storing programs or instructions and a second storage area for storing data. The first storage area can store the operating system, application programs or instructions required for at least one function (such as sound playback, image playback, etc.). Furthermore, memory can include volatile memory or non-volatile memory, or both. Non-volatile memory can be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. Volatile memory can be random access memory (RAM), static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDRSDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous linked dynamic random access memory (Synchlink DRAM, SLDRAM), and direct memory bus RAM (DRRAM).
[0138] The processor may include one or more processing units; optionally, the processor integrates an application processor and a modem processor, wherein the application processor mainly handles operations related to the operating system, user interface, and applications, while the modem processor mainly handles wireless communication signals, such as a baseband processor. It is understood that the aforementioned modem processor may also not be integrated into the processor.
[0139] This application also provides a computer-readable storage medium for storing computer-executable instructions. When these computer-executable instructions are executed by a processor, they implement the various processes of the above-described dialogue summary generation method embodiments and achieve the same technical effects. To avoid repetition, these will not be described again here.
[0140] The processor is the processor in the electronic device described in the above embodiments. The readable storage medium includes computer-readable storage media, such as computer read-only memory (ROM), random access memory (RAM), magnetic disk, or optical disk.
[0141] This application also provides a computer program product, which includes a computer program. When the computer program is executed by a processor, it implements the various processes of the above-described dialogue summary generation method embodiments and can achieve the same technical effect. To avoid repetition, it will not be described again here.
[0142] This application embodiment also provides a chip, which includes a processor and a communication interface. The communication interface is coupled to the processor. The processor is used to run programs or instructions to implement the various processes of the above-described dialogue summary generation method embodiment and can achieve the same technical effect. To avoid repetition, it will not be described again here.
[0143] It should be understood that the chip mentioned in the embodiments of this application may also be referred to as a system-on-a-chip, system chip, chip system, or system-on-a-chip, etc.
[0144] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one…" does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element. Furthermore, it should be noted that the scope of the methods and apparatuses in the embodiments of this application is not limited to performing functions in the order shown or discussed, but may also include performing functions substantially simultaneously or in the reverse order, depending on the functions involved. For example, the described methods may be performed in a different order than described, and various steps may be added, omitted, or combined. Additionally, features described with reference to certain examples may be combined in other examples.
[0145] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal (which may be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in the various embodiments of this application.
[0146] The embodiments of this application have been described above with reference to the accompanying drawings. However, this application is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of this application without departing from the spirit and scope of the claims, and all of these forms are within the protection scope of this application.< / summary> < / summary> < / entity> < / entity>
Claims
1. A dialogue summary generation method, characterized in that, include: Obtain the text of the conversation between the user and the agent; The dialogue text is input into a pre-trained scene type classification model, and the target scene type corresponding to the dialogue text is output. Based on the target scene type, a target summary template matching the target scene type is determined; The dialogue text, the target summary template, and the first dialogue summary generation task prompt are input into the dialogue summary generation model. The first dialogue summary generation task prompt guides the dialogue summary generation model to output a first dialogue summary that corresponds to the dialogue text and conforms to the target summary template. The dialogue summary generation model is obtained by optimizing the first large language model based on reinforcement learning fine-tuning.
2. The method according to claim 1, characterized in that, The training process of the dialogue summary generation model includes: Obtain the first training dataset; the first training dataset includes multiple sets of first training data, each set of first training data includes a first sample dialogue text, as well as a sample dialogue summary and sample entity information corresponding to the first sample dialogue text; Based on the first subset within the first training dataset, the first large language model is fine-tuned using a multi-task joint training method to obtain the second large language model; the multi-task joint training method includes training the first large language model through entity information extraction tasks and dialogue summary generation tasks. Based on the second subset within the first training dataset, the second large language model is fine-tuned using a reinforcement learning algorithm to maximize the reward value of the target reward function, thereby obtaining the dialogue summary generation model.
3. The method according to claim 2, characterized in that, The second language model is obtained by fine-tuning the first large language model using a multi-task joint training method based on a first subset of the first training dataset, including: For each set of first training data in the first subset, a first instruction fine-tuning sample and a second instruction fine-tuning sample are constructed based on the first training data. The first instruction fine-tuning sample includes the first training data and entity information extraction task prompts. The second instruction fine-tuning sample includes the first sample dialogue text, the sample dialogue summary, the sample summary template corresponding to the first sample dialogue text, and the second dialogue summary generation task prompts. The first instruction fine-tuning sample is input into the first large language model, and the first large language model is guided to output the first predicted entity information corresponding to the first sample dialogue text through the entity information extraction task prompt words; the first loss result of the first large language model under the entity information extraction task is determined according to the sample entity information, the first predicted entity information and the first loss function. The second instruction fine-tuning sample is input into the first large language model, and the first large language model is guided to output the first predicted dialogue summary corresponding to the first sample dialogue text by the second dialogue summary generation task prompt words; the second loss result of the first large language model under the dialogue summary generation task is determined according to the sample dialogue summary, the first predicted dialogue summary and the second loss function. Based on the first loss result and the second loss result of the first large language model, the model parameters of the first large language model are updated to obtain the second large language model.
4. The method according to claim 2, characterized in that, The step of fine-tuning the second language model using a reinforcement learning algorithm based on a second subset of the first training dataset to maximize the reward result of the target reward function, thereby obtaining the dialogue summary generation model, includes: For each set of first training data in the second subset, the first training data, the sample summary template corresponding to the first sample dialogue text, and the first task prompt word are input into the second large language model. The second large language model is guided by the first task prompt word to output the second predicted entity information and the second predicted dialogue summary corresponding to the first sample dialogue text. Based on the sample entity information, the second predicted entity information, and a preset entity reward function, the entity reward result of the second large language model is determined; and based on the sample dialogue summary, the second predicted dialogue summary, and a preset edit distance reward function, the edit distance reward result of the second large language model is determined; and based on the sample summary template, the second predicted dialogue summary, and a preset template format reward function, the template format reward result of the second large language model is determined. Based on the entity reward result, the edit distance reward result, the template format reward result, and the KL divergence of the second largest language model, determine the reward result of the objective reward function of the second largest language model; The PPO algorithm is optimized based on the near-end strategy to update the model parameters of the second language model until the reward result of the target reward function reaches the set condition or the number of updates reaches the set number, thus obtaining the dialogue summary generation model.
5. The method according to claim 1, characterized in that, The method further includes: Obtain a second dialogue summary; the second dialogue summary is generated by the agent after editing the first dialogue summary; The first dialogue summary and the second dialogue summary are compared and analyzed to generate a multi-dimensional evaluation result of the dialogue summary generation model; the multi-dimensional evaluation result includes several of the following: incorrectly predicted entity names, unpredicted entity names, and incorrectly predicted scenario types; Based on the multi-dimensional evaluation results, targeted training data for optimizing the dialogue summary generation model is constructed through an intelligent agent.
6. The method according to claim 1, characterized in that, The training process of the scene type classification model includes: Obtain a second training dataset; the second training dataset includes multiple sets of second training data, each set of second training data includes second sample dialogue text and the second sample scene type corresponding to the second sample dialogue text; For each set of the second training data, the core dialogue question corresponding to the second sample dialogue text is extracted using a large model; Based on the second sample dialogue text, the second sample scene type, and the core dialogue question, construct the third training data; The third training data is input into the BERT model to be trained, and the predicted scene type corresponding to the second sample dialogue text is output. Based on the second sample scenario type, the predicted scenario type, and the preset loss function, the third loss result of the BERT model is determined; Based on the third loss result, the model parameters of the BERT model are updated to obtain the scene type classification model.
7. A dialogue summarization generation apparatus, characterized in that, include: The first acquisition module is used to acquire the dialogue text between the user and the agent; The scene type classification processing module is used to input the dialogue text into a pre-trained scene type classification model and output the target scene type corresponding to the dialogue text. The determination module is used to determine a target summary template that matches the target scene type based on the target scene type. The dialogue summarization module is used to input the dialogue text, the target summarization template, and the first dialogue summarization task prompt into the dialogue summarization model. The first dialogue summarization task prompt guides the dialogue summarization model to output a first dialogue summary that corresponds to the dialogue text and conforms to the target summarization template. The dialogue summarization model is obtained by optimizing the first large language model based on reinforcement learning fine-tuning.
8. An electronic device, characterized in that, include: processor; as well as A memory configured to store computer-executable instructions configured to be executed by the processor to implement the dialogue summary generation method as described in any one of claims 1-6.
9. A computer-readable storage medium, characterized in that, The computer-readable storage medium is used to store computer-executable instructions, which, when executed by a processor, implement the dialogue summary generation method as described in any one of claims 1-6.
10. A computer program product, characterized in that, It includes a computer program that, when executed by a processor, implements the dialogue summary generation method as described in any one of claims 1-6.