A method for training an abstract generation model, an abstract generation method, an apparatus, and a device
Patent Information
- Application Number
- CN202410020385.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-01-05
- Publication Date
- 2026-09-08
- Estimated Expiration
- 2044-01-05
AI Technical Summary
但是,目前的自然语言处理模型通常较难胜任特定场景下细粒度摘要的生成任务,如生成包含通话记录中提及的核心事项和待办事项的摘要
[0038] This application provides a training scheme for a summary generation model. After obtaining sample call records with labeled call subjects, a natural language processing model processes the sample call records according to guiding text to obtain a summary with any call subject as the subject and containing multiple summary attributes. Then, the summary generation model is trained using the sample call records and their summaries, enabling the model to generate fine-grained call record summaries. That is, the summary generation model can generate summaries containing multiple summary attributes with any call subject in the call record as the subject, thereby ensuring that the summary contains the key content of the call record and improving the comprehensiveness and accuracy of the generated summary.
Smart Images

Figure CN117851584B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and in particular to a method for training a summary generation model, a summary generation method, an apparatus, and a device. Background Technology
[0002] With the development of deep learning technology, summarization has gradually become an important research direction in the field of natural language processing. The goal of summarization is to extract key information from given raw text and present it to the user clearly and concisely. Among related technologies, many deep learning-based natural language processing models perform well on general summarization tasks. However, current natural language processing models often struggle with generating fine-grained summaries for specific scenarios, such as generating summaries that include core matters and to-do items mentioned in call logs. Therefore, there is an urgent need for a summarization model capable of accurately generating fine-grained call log summaries. Summary of the Invention
[0003] This application provides a training method, a method for generating a summary model, an apparatus, and a device for doing so. The trained summary model can generate fine-grained call record summaries, ensuring that the call record summaries contain key content from the call records, thus improving the comprehensiveness and accuracy of the generated summaries. The technical solution is as follows:
[0004] According to one aspect of the embodiments of this application, a method for training a summary generation model is provided, the method comprising:
[0005] Obtain sample call records, wherein the sample call records are marked with call recipients and the sample call records are the call content of the call recipients;
[0006] The sample call record is processed by a natural language processing model based on guiding text to obtain a first summary. The natural language processing model is used to process the input text data to obtain a summary of the text data. The guiding text is used to guide the natural language processing model to generate a summary of the sample call record in the form of a first-person subject. The summary contains multiple summary attributes, and the summary attributes are used to represent at least one key content in the sample call record.
[0007] Based on the sample call records and the first summary, a summary generation model is trained so that the summary generation model can process the call records to generate a summary of the call records.
[0008] According to another aspect of the embodiments of this application, a method for generating a digest is provided, the method comprising:
[0009] In response to the end of the call, the call log and thought chain of the call are input into the summary generation model. The call log is marked with the call subject and is the content of the call of the call subject. The thought chain is used to indicate multiple preset questions. The multiple preset questions are used to extract keywords and key sentences in the call log. The summary generation model is used to generate a summary of the call log in the form of a first-person subject. The summary contains multiple summary attributes. The summary attributes are used to represent at least one key content in the call log.
[0010] The call record is processed using the summary generation model and based on the thought chain to obtain the answer text for the multiple preset questions. The answer text includes at least one of the keywords and key phrases.
[0011] The summary generation model generates a summary of the call record based on the call record and the answer text of the multiple preset questions.
[0012] According to another aspect of the embodiments of this application, a training apparatus for a summary generation model is provided, the apparatus comprising:
[0013] The first acquisition module is used to acquire sample call records, wherein the sample call records are marked with call objects and the sample call records are the call content of the call objects;
[0014] The first generation module is used to process the sample call record using a natural language processing model based on guiding text to obtain a first summary. The natural language processing model is used to process the input text data to obtain a summary of the text data. The guiding text is used to guide the natural language processing model to generate a summary of the sample call record in the form of a first-person subject. The summary contains multiple summary attributes, and the summary attributes are used to represent at least one key content in the sample call record.
[0015] The training module is used to train the summary generation model based on the sample call records and the first summary, so that the summary generation model can process the call records to generate a summary of the call records.
[0016] In some embodiments, the training module is configured to process the sample call records based on the summary generation model to obtain a second summary; determine the training loss of the summary generation model based on the first summary and the second summary; and train the summary generation model based on the training loss.
[0017] In some embodiments, the apparatus further includes:
[0018] The second generation module is used to generate the guidance text, which includes a first rule and a second rule. The first rule is used to define multiple summary attributes that need to be included in the summary, and the second rule is used to define the generation rules and display format of the multiple summary attributes.
[0019] An input module is used to input the guiding text and the sample call record into the natural language processing model.
[0020] In some embodiments, the input module is further configured to obtain a summary sample, wherein the summary sample is a call record summary that conforms to the first rule and the second rule; and input the summary sample, the guiding text, and the sample call record into the natural language processing model.
[0021] In some embodiments, the apparatus further includes:
[0022] The display module is used to display a first prompt message when the first summary is obtained. The first prompt message is used to prompt whether the format and content of the first summary should be corrected.
[0023] The second acquisition module is used to acquire the corrected first summary in response to the confirmation operation of the first prompt information.
[0024] In some embodiments, the apparatus further includes:
[0025] A pre-training module is used to pre-train the summary generation model based on a pre-training dataset, which includes terms and nouns from multiple domains. The pre-training is used to enable the summary generation model to understand the terms and nouns from the multiple domains in order to generate summaries of call records from the multiple domains.
[0026] According to another aspect of the embodiments of this application, a digest generation apparatus is provided, the apparatus comprising:
[0027] An input module is used to input the call log and thought chain of the call into a summary generation model in response to the end of the call. The call log is marked with the call subject and is the content of the call of the call subject. The thought chain is used to indicate multiple preset questions. The multiple preset questions are used to extract keywords and key sentences from the call log. The summary generation model is used to generate a summary of the call log in the form of a first-person subject. The summary contains multiple summary attributes, and the summary attributes are used to represent at least one key content in the call log.
[0028] The first generation module is used to process the call record based on the thought chain through the summary generation model to obtain the answer text of the multiple preset questions, wherein the answer text includes at least one of the keywords and the key sentences;
[0029] The second generation module is used to generate a summary of the call record based on the call record and the answer text of the multiple preset questions through the summary generation model.
[0030] In some embodiments, the input module is configured to, in response to the end of a call, convert the call audio to text to obtain the call content; based on the caller and the called party, label the call content with the call objects to obtain the call record; and input the call record and the thought chain into the summary generation model.
[0031] In some embodiments, the second generation module is configured to input the call record, the answer text of the plurality of preset questions, and the preset task text into the summary generation model, so as to generate a summary of the call record by the summary generation model according to the task text, wherein the task text is used to instruct the summary generation model to generate a summary of the call record based on the call record and the answer text of the plurality of preset questions.
[0032] In some embodiments, the apparatus further includes:
[0033] The display module is configured to display a second prompt message when a summary of the call record is obtained, the second prompt message being used to prompt whether to display the summary of the call record; and to display the summary of the call record in response to a confirmation operation of the second prompt message.
[0034] According to another aspect of the embodiments of this application, a computer device is provided, the computer device including a processor and a memory; the memory stores at least one piece of program code, the at least one piece of program code being executed by the processor to implement the training method of the summary generation model as described above, or to implement the summary generation method as described above.
[0035] According to another aspect of the embodiments of this application, a chip is provided, the chip including programmable logic circuits and / or program instructions, which, when the chip is run on a computer device, are used to implement the training method of the summary generation model described above, or to implement the summary generation method as described above.
[0036] According to another aspect of the embodiments of this application, a computer-readable storage medium is provided, the storage medium storing at least one piece of program code, the at least one piece of program code being executed by a processor to implement the training method of the summarization generation model as described above, or to implement the summarization generation method as described above.
[0037] According to another aspect of the embodiments of this application, a computer program product is provided, which stores at least one piece of program code, the at least one piece of program code being executed by a processor to implement the training method of the summary generation model described above, or to implement the summary generation method as described above.
[0038] This application provides a training scheme for a summary generation model. After obtaining sample call records with labeled call subjects, a natural language processing model processes the sample call records according to guiding text to obtain a summary with any call subject as the subject and containing multiple summary attributes. Then, the summary generation model is trained using the sample call records and their summaries, enabling the model to generate fine-grained call record summaries. That is, the summary generation model can generate summaries containing multiple summary attributes with any call subject in the call record as the subject, thereby ensuring that the summary contains the key content of the call record and improving the comprehensiveness and accuracy of the generated summary. Attached Figure Description
[0039] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0040] Figure 1 This is a schematic diagram of the implementation environment of a training method for a summary generation model provided in an embodiment of this application;
[0041] Figure 2 This is a flowchart illustrating a training method for a summary generation model provided in an embodiment of this application;
[0042] Figure 3 This is a flowchart of another training method for a summary generation model provided in an embodiment of this application;
[0043] Figure 4 This is a flowchart of an abstract generation method provided in an embodiment of this application;
[0044] Figure 5This is a schematic diagram of the structure of a training device for a summary generation model provided in an embodiment of this application;
[0045] Figure 6 This is a schematic diagram of the structure of a training device for another summary generation model provided in an embodiment of this application;
[0046] Figure 7 This is a schematic diagram of the structure of an abstract generation device provided in an embodiment of this application;
[0047] Figure 8 This is a schematic diagram of another abstract generation apparatus provided in an embodiment of this application;
[0048] Figure 9 This is a structural block diagram of a terminal provided in an embodiment of this application. Detailed Implementation
[0049] To make the objectives, technical solutions, and advantages of this application clearer, the embodiments of this application will be described in further detail below with reference to the accompanying drawings.
[0050] In this article, "at least one" refers to one or more, and "multiple" refers to two or more. "And / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, or B alone. The character " / " generally indicates that the preceding and following related objects have an "or" relationship.
[0051] It should be noted that all information (including but not limited to user device information, user personal information, etc.), data (including but not limited to data used for analysis, stored data, displayed data, etc.), and signals involved in this application have been authorized by the user or fully authorized by all parties, and the collection, use, and processing of related data must comply with the relevant laws, regulations, and standards of the relevant countries and regions. For example, call records and call content involved in this application were obtained with full authorization.
[0052] The method for training the summary generation model provided in this application can be executed by a computer device. In some embodiments, the computer device is a terminal or a server. The implementation environment of the method for training the summary generation model provided in this application is described below.
[0053] Figure 1 This is a schematic diagram illustrating the implementation environment of a training method for a summary generation model provided in this application. See also... Figure 1 The implementation environment includes a terminal 101 and a server 102. The terminal 101 can be directly or indirectly connected to the server 102 via wired or wireless communication.
[0054] In some embodiments, terminal 101 can be various types of terminals such as smartphones, smartwatches, desktop computers, laptops, and tablets. Terminal 101 has a communication application installed. Users can use the communication application to control terminal 101 to establish voice or video calls with other terminals. Terminal 101 also has a summary generation model deployed. The summary generation model is used to generate a summary of the call log so that users can view the key content in the call log. The summary of the call log may include a summary of the call log, key matters mentioned in the call log, and to-do items. After user authorization, terminal 101 can input the call content into the summary generation model after the voice or video call ends, and the summary generation model will generate a summary of the voice or video call. This summary generation model is associated with server 102, which provides services such as training, updating, and on-device deployment of the summary generation model.
[0055] In some embodiments, server 102 can be an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms.
[0056] In some embodiments, server 102 undertakes the main computing work and terminal 101 undertakes the secondary computing work; or, server 102 undertakes the secondary computing work and terminal 101 undertakes the main computing work; or, server 102 and terminal 101 collaborate on computing using a distributed computing architecture.
[0057] Figure 2 This is a flowchart illustrating a training method for a summary generation model provided in an embodiment of this application. The method is executed by a computer device; see [link to relevant documentation]. Figure 2 The method includes:
[0058] 201. Computer equipment acquires sample call records. The sample call records are marked with the call recipients and contain the call content of the call recipients.
[0059] In this embodiment, the sample call logs include the call content of the call parties. The call content in the sample call logs is labeled with the call parties. The sample call logs are used to train a summary generation model, enabling the model to generate a summary of the call logs with the call parties as the subject.
[0060] Computer devices can acquire sample call records in various ways. For example, a computer device can retrieve sample call records from a local database. This local database stores the computer device's historical call records. Accordingly, after ending a call, the computer device stores the call record labeled with the caller in the local database. Alternatively, the computer device can acquire sample call records uploaded by multiple terminals. Accordingly, a terminal can upload call records labeled with the caller to the computer device with user authorization.
[0061] 202. The computer device processes the sample call record using a natural language processing model based on the guiding text to obtain a first summary. The natural language processing model is used to process the input text data to obtain a summary of the text data. The guiding text is used to guide the natural language processing model to generate a summary of the sample call record in the form of a first-person subject. The summary contains multiple summary attributes, which are used to represent at least one key content in the sample call record.
[0062] In this embodiment, the computer device inputs a sample call log and guiding text into a natural language processing (NLP) model, enabling the NLP model to generate a summary of the sample call log based on the guiding text. For ease of description, the summary generated by the NLP model is referred to as the first summary. The guiding text is used to guide the NLP model in generating a first summary that meets the requirements. The guiding text guides the NLP model to generate the first summary in the form of a first-person subject, and the generated first summary needs to contain multiple summary attributes. These summary attributes are part of the summary and represent at least one key element of the sample call log.
[0063] For example, the first summary includes three summary attributes: Summary, Core Items, and To-Do Items. The Summary represents a summary of the topics or themes discussed in the sample call log. The Core Items represent at least one core event mentioned in the sample call log. The To-Do Items represent at least one pending event mentioned in the sample call log. The Summary, Core Items, and To-Do Items in the first summary are all presented in the first person as subjects.
[0064] 203. The computer equipment trains the summary generation model based on sample call records and a first summary, so that the summary generation model can process the call records to generate a summary of the call records.
[0065] In this embodiment, the computer device trains the summary generation model using sample call records and a first summary as the training sample set, resulting in a trained summary generation model. Since the first summary meets preset requirements, training the summary generation model with the aforementioned training sample set enables the model to gradually acquire the ability to generate summaries that meet these requirements during the training process. Processing any call record with the trained summary generation model yields a summary of that call record generated by the model using a first-person subject, thus ensuring the comprehensiveness and accuracy of the summary.
[0066] By generating call log summaries using a summary generation model, users can gain a clear understanding of the call log summary, key events mentioned, and to-do items without having to listen to the recording again. This improves the efficiency of users in obtaining key information from call logs.
[0067] This application provides a method for training a summary generation model. After obtaining sample call records with labeled call subjects, a natural language processing model processes the sample call records according to guiding text to obtain a summary with any call subject as the subject and containing multiple summary attributes. Then, the summary generation model is trained using the sample call records and their summaries, enabling the model to generate fine-grained call record summaries. That is, the summary generation model can generate summaries with multiple summary attributes, using any call subject in the call record as the subject, thus ensuring that the summary contains the key content of the call record and improving the comprehensiveness and accuracy of the generated summary.
[0068] Figure 3 This is a flowchart illustrating another training method for a summary generation model provided in this application embodiment. The method is executed by a computer device; see [link to relevant documentation]. Figure 3 The method includes:
[0069] 301. The computer device pre-trains a summary generation model based on a pre-trained dataset, which includes terms and nouns from multiple domains. The pre-training is used to enable the summary generation model to understand the terms and nouns from multiple domains in order to generate summaries of call records from multiple domains.
[0070] In this embodiment, the pre-training dataset is used to pre-train the summarization generation model. The pre-training dataset includes terminology from multiple domains. The summarization generation model is a natural language processing model used to generate summaries of call logs, such as a Large Language Model (LLM) or LLaMA (Large Language Model Meta Artificial Intelligence). Natural language processing models aim to understand and generate natural language, handling various tasks in the field of natural language processing, such as text summarization and text abstraction.
[0071] To ensure the summary generation model can accurately generate summaries of call records across multiple domains, the computer device first pre-trains the model using a pre-training dataset. This pre-training allows the model to acquire prior knowledge and common sense across various domains, thereby improving its performance in call record summary generation tasks across different domains.
[0072] The pre-training dataset may include terminology from various industries, such as the internet, mobile phone, automotive, catering, education, and tourism sectors. It may also include everyday terminology, such as those related to entertainment, social interaction, emotions, shopping, fashion, and cooking. Furthermore, the pre-training dataset may include more specialized terminology, such as academic, medical, and engineering terms; however, this embodiment does not impose any limitations on this.
[0073] 302. Computer equipment obtains sample call records. The sample call records are marked with the call recipients and contain the call content of the call recipients.
[0074] In this embodiment, the sample call logs are call logs used for formal training of the summary generation model. The sample call logs include the content of the calls between the callers. Each call content in the sample call logs is labeled with its corresponding caller. The call content in the sample call logs can be multiple call texts arranged chronologically, with each call text labeled with its corresponding caller.
[0075] Computer devices can acquire sample call logs in various ways. For example, when the computer device is a terminal, it can retrieve call logs from a local database. This local database stores multiple historical call logs of the terminal. Accordingly, with user authorization, the terminal can store the call logs of a voice call or video call in the local database after ending the call. When the computer device is a server, the server can acquire sample call logs uploaded by multiple terminals. Accordingly, with user authorization, the terminal can also upload multiple historical call logs of the terminal to the server.
[0076] In some embodiments, the terminal can generate a call log after the call ends. In response to the end of a voice or video call, the terminal converts the call audio to text to obtain multiple call texts. Then, based on the caller and callee, the terminal sequentially marks the corresponding call recipient before each call text, resulting in the following call log: My side: Call text 1. Other party: Call text 2. My side: Call text 3. Other party: Call text 4… The terminal can then store the generated call logs in a local database, or it can upload the call logs to a server to provide sample call logs to the server.
[0077] 303. The computer device generates guidance text, which includes a first rule and a second rule. The first rule is used to define the multiple summary attributes that need to be included in the summary, and the second rule is used to define the generation rules and display format of the multiple summary attributes.
[0078] In this embodiment of the application, in order to enrich the training sample set of the summary generation model, after the computer device obtains the sample call records, the computer device processes the sample call records through a natural language processing model to obtain a summary of the sample call records. For ease of description, the summary generated by the natural language processing model will be referred to as the first summary below.
[0079] The natural language processing model can be a pre-trained ChatGPT (Chat Generative Pre-trained Transformer), BERT (Bidirectional Encoder Representations from Transformers), PaLM (Pathways Language Model), or LLaMA (Large Language Model Meta Artificial Intelligence), and this application does not limit it.
[0080] Before processing the sample call log using a natural language processing (NLP) model, the computer device first generates guidance text. Guidance text refers to text instructions input into the NLP model to guide it in generating specific outputs. Guidance text typically describes information, answers, or text that the user wants to obtain from the NLP model. Guidance text can be a prompt input into the NLP model. In this embodiment, the guidance text guides the NLP model to generate a summary of the sample call log in a first-person subject format, and this summary needs to contain multiple summary attributes. These summary attributes are part of the summary. The summary attributes represent at least one key element of the sample call log.
[0081] For example, the summary attribute includes a summary, core items, and to-do items. A summary represents a general overview of the topics or themes discussed in the sample call log. Core items represent at least one core event mentioned in the sample call log. To-do items represent at least one pending event mentioned in the sample call log.
[0082] During the generation of the guidance text, the computer device can write a first rule and a second rule into the guidance text. The first rule defines the multiple summary attributes that the summary must include. For example, for the three summary attributes mentioned above, the first rule defines "summary" and "core matters" as required summary attributes, while "to-do items" is optional. Therefore, the first summary generated by the natural language processing model must contain a summary and core matters; it may or may not contain to-do items. The second rule defines the generation rules and display format for each summary attribute. By specifying multiple rules for generating the summary in the guidance text, the natural language processing model can generate the summary according to these rules, ensuring that the summary meets the preset requirements in both format and content.
[0083] Table 1 below shows the generation rules and display format for each summary attribute defined by the second rule.
[0084] Table 1
[0085]
[0086]
[0087] It should be noted that Table 1 above presents an example of a second rule in tabular form. In some embodiments, those skilled in the art can modify the generation rules and display format defined in the second rule according to actual needs. Furthermore, this application embodiment uses a tabular format to present the second rule in order to clearly and intuitively illustrate the generation rules and display format defined in the second rule. During the generation of the guidance text, the computer device writes the first and second rules into the guidance text in text form so that the text-based guidance text can be subsequently input into the natural language processing model.
[0088] For example, the guidance text obtained by the computer device through step 303 can be: "Generate a summary of the above call record in the form of a first-person subject. The summary should include a summary, core matters, and to-do items. The generation rules and display formats of the summary, core matters, and to-do items are xxxx respectively."
[0089] It should be noted that this embodiment of the application illustrates the example where the computer device executes step 302 first, followed by step 303. In some embodiments, the computer device may execute step 303 first, followed by step 302; that is, the computer device generates the guiding text before acquiring the sample call record. This embodiment of the application does not impose any restrictions on the order in which the computer device executes steps 302 and 303.
[0090] 304. The computer device inputs the guiding text and sample call records into the natural language processing model, so that the natural language processing model processes the sample call records based on the guiding text to obtain a first summary.
[0091] In this embodiment, the computer device inputs the acquired sample call records and the generated guidance text into a natural language processing model. Then, the computer device processes the sample call records according to the guidance text through the natural language processing model to obtain a first summary that conforms to various rules defined in the guidance text.
[0092] In some embodiments, during the process of generating the first summary with a first-person subject, the natural language processing (NLP) model can also perform semantic analysis on the call content in the sample call record to obtain the identity information of the two parties in the call. Then, based on the identity information of the two parties, the NLP model generates the first summary with a first-person subject. For example, if the NLP model determines that the caller is a job seeker and the other party is an HR employee, the first summary generated by the NLP model could be: 1. I inquired about the xx position. 2. HR explained the job content of the xx position. 3. I expressed my doubts about the job content.
[0093] In some embodiments, the computer device can also input a summary sample into the natural language processing model. The computer device obtains a summary sample, which is a call record summary conforming to the first and second rules described above. Then, the computer device inputs the summary sample, the guiding text, and the sample call record into the natural language processing model. By inputting a summary sample conforming to multiple rules defined in the guiding text into the natural language processing model, the model can further understand the multiple rules in the guiding text based on the summary sample, thereby outputting a first summary that better conforms to the aforementioned multiple rules, improving the accuracy of the first summary.
[0094] For example, the following are sample summaries that conform to the first and second rules above.
[0095] Summary: This phone call was to communicate with the HR staff of Company X regarding information related to the xx position.
[0096] Key issues:
[0097] 1. I inquired about the job responsibilities, work experience requirements, salary range, and work location for the xx position.
[0098] 2. The HR staff explained the job content and job level range of the xx position.
[0099] 3. I have questions about the job responsibilities of the xx position.
[0100] 4. I inquired about the status of my resume updates.
[0101] 5. I expressed my hope to find job opportunities in the fields of voice or operating systems.
[0102] To-do list:
[0103] 1. The HR staff needs to update my resume information.
[0104] 2. The HR staff needs to relay my needs and intentions to the interviewer.
[0105] 3. If the HR staff can find a suitable position within this week, they will contact me within this week.
[0106] As can be seen from the above summary example, the summary example includes three summary attributes: summary, core items, and to-do items. The core items section displays multiple core events from the call log, and the to-do items section displays multiple to-do events mentioned in the call log.
[0107] In some embodiments, during the generation of core items in the first summary, the natural language processing model can sequentially display multiple core events in chronological order, using multiple sub-points. For example, the five core events in the above summary example are arranged in chronological order. Alternatively, the natural language processing model can also display multiple core events sequentially in descending order of priority, as described in this embodiment. The priority of a core event is positively correlated with its importance. Furthermore, the natural language processing model can also determine the display order of multiple to-do items in a similar manner, which will not be elaborated further here.
[0108] 305. Upon receiving the first summary, the computer device displays a first prompt message, which prompts whether the format and content of the first summary should be corrected.
[0109] In this embodiment, after the computer device obtains the first summary output by the natural language processing model, the computer device can display the first summary and a first prompt message. The first prompt message asks the user whether the format and content of the first summary need to be corrected. By reviewing the first summary displayed by the computer device, the user can determine whether the format and content of the first summary are correct. If the format or content of the first summary is incorrect, the user can manually correct it. For example, the user could be a technician who trained the summary generation model.
[0110] 306. In response to the confirmation operation of the first prompt information, the computer device obtains the corrected first summary.
[0111] In this embodiment, a technician confirms the correction of the first summary by performing a confirmation operation on the first prompt information. The confirmation operation can be a triggering operation of the confirmation control displayed in the first prompt information, or a clicking or long-pressing operation on the first prompt information; this embodiment does not limit this. In response to the confirmation operation on the first prompt information, the computer device displays an editing page for the first summary. The technician can edit the format or content of the first summary on the editing page to correct it. In response to the technician's submission operation of the first summary on the editing page, the computer device obtains the manually corrected first summary from the editing page. By supporting manual correction of the first summary by technicians, the correctness of the format and content of the first summary can be ensured, avoiding the situation where the summary generation model is trained with an incorrect first summary, thereby ensuring the training effect of the summary generation model.
[0112] For example, if there are no core events in the core items of the first summary, technicians can manually enter multiple core events under the core items section on the editing page, based on the sample call log. Alternatively, if there are obvious grammatical errors in the summary of the first summary, technicians can manually summarize the sample call log and manually enter the summary text under the summary section on the editing page.
[0113] 307. The computer device trains a summary generation model based on sample call records and a first summary, so that the summary generation model can process the call records to generate a summary of the call records.
[0114] In this embodiment, after obtaining the first summary, the computer device uses the sample call records and the first summary as a training sample set to formally train the summary generation model. Since the content and format of the first summary are correct, and the first summary conforms to multiple rules defined in the guiding text, formally training the summary generation model using the aforementioned training sample set enables the model to gradually acquire the ability to generate summaries that meet preset requirements during the training process.
[0115] After the aforementioned pre-training and formal training, the summary generation model possesses the ability to generate summaries of call logs across multiple domains. Processing call logs from any domain using the summary generation model yields summaries that meet preset requirements in both content and format. These summaries possess multiple summary attributes specified in the guiding text in terms of format; and in terms of content, they accurately and comprehensively reflect the key information in the call log.
[0116] In some embodiments, the computer device can train a summarization generation model based on a training loss. First, the computer device inputs sample call records into the summarization generation model to process the records and obtain a summary of the call records. For ease of description, the summary generated by the model is referred to as the second summary. Then, the computer device determines the training loss of the summarization generation model based on the first and second summaries. The training loss is positively correlated with the difference between the first and second summaries. The computer device can determine the difference between the first and second summaries as the distance between the feature vectors of the first and second summaries. Then, the computer device updates the model parameters of the summarization generation model based on the training loss to reduce the training loss. If the updated model meets the training termination condition, such as reaching the target number of training iterations or the training loss being within the target range, the updated model is considered a completed training model. If the updated model does not meet the training termination condition, the parameters are updated continuously until the model meets the training termination condition, resulting in a completed training model. By training the summary generation model using the difference between the first and second summaries, which is the training loss, the summary generation model can gradually acquire the ability to generate call record summaries that meet preset requirements, thus obtaining a summary generation model that can accurately generate quasi-fine-grained call record summaries.
[0117] This application provides a method for training a summary generation model. After obtaining sample call records with labeled call subjects, a natural language processing model processes the sample call records according to guiding text to obtain a summary with any call subject as the subject and containing multiple summary attributes. Then, the summary generation model is trained using the sample call records and their summaries, enabling the model to generate fine-grained call record summaries. That is, the summary generation model can generate summaries with multiple summary attributes, using any call subject in the call record as the subject, thus ensuring that the summary contains the key content of the call record and improving the comprehensiveness and accuracy of the generated summary.
[0118] The above embodiments mainly describe the process of training a summary generation model using a computer device. After the computer device completes the training of the summary generation model, when the computer device is a terminal, the terminal can directly use the trained summary generation model on the device side to generate a summary of call records. When the computer device is a server, the server can deploy the summary generation model on the terminal, so that the terminal can use the summary generation model on the device side. The following embodiments illustrate the process of a terminal using the device-side summary generation model to generate a summary of call records.
[0119] Figure 4 This is a flowchart of a summary generation method provided in an embodiment of this application. The method is executed by a terminal; see [link to relevant documentation]. Figure 4 The method includes:
[0120] 401. In response to the end of the call, the terminal inputs the call log and thought chain into the summary generation model. The call log is marked with the call subject and contains the call content of the call subject. The thought chain is used to indicate multiple preset questions. The multiple preset questions are used to extract keywords and key sentences from the call log. The summary generation model is used to generate a summary of the call log in the form of a first-person subject. The summary contains multiple summary attributes, which are used to represent at least one key content in the call log.
[0121] In this embodiment, the terminal is a device with calling capabilities, such as a mobile phone, smartwatch, tablet computer, or laptop computer. The terminal has communication applications installed, such as telephones or social applications that support audio and video calls. Users can use these communication applications to establish audio and video calls with other terminals. After a user initiates an audio or video call through the terminal, in response to the end of the call, the terminal retrieves the call log. The call log contains the call content annotated with the call participants. Then, the terminal inputs the call log and a preset thought chain into a summary generation model on the client side. The thought chain includes multiple preset questions. These preset questions are used to extract key information from the call log, such as keywords related to important entities and dates, or key statements about events mentioned and their outcomes.
[0122] For example, the thought process might include the following four questions: 1. What are the important entities in the conversation? 2. What are the important dates in the conversation? 3. What were discussed in the conversation? 4. What were the results of these discussions?
[0123] In some embodiments, the terminal can generate a call log after the call ends. In response to the end of the call, the terminal converts the call audio to text to obtain the call content. The call content includes multiple text messages. Then, the terminal labels the multiple text messages in the call content with call subjects based on the caller and callee, obtaining the call log. The terminal then inputs the call log and thought chain into a summary generation model. By converting the call audio to text to obtain the call content and labeling the call content with call subjects, the summary generation model can analyze the identity information of both parties based on the labeled call content, thereby accurately generating a summary with a first-person subject.
[0124] 402. The terminal processes the call record through a summary generation model based on the thought chain to obtain answer texts for multiple preset questions. The answer texts include at least one of keywords and key phrases.
[0125] In this embodiment, the summary generation model is a natural language processing model. Therefore, it can generate fine-grained call record summaries not only in the first-person subject format, but also perform various tasks in the natural language processing domain, such as text summarization, text translation, sentiment analysis, and key information extraction. Thus, after the terminal inputs the call record and thought chain into the summary generation model, the model understands multiple preset questions in the thought chain and sequentially answers these questions by understanding, analyzing, and extracting key information from the call record, thereby obtaining the answer text for each preset question. The answer text includes at least one keyword or key phrase extracted from the call record by the summary generation model. Extracting key information from the call record through multiple preset questions in the thought chain not only reduces the difficulty of key information extraction for the summary generation model but also ensures the controllability of the extracted key information.
[0126] For example, the answer texts to the four preset questions obtained through the summary generation model could be as follows: 1. The important entities in this call are: landlord and tenant. 2. The important date in this call is: the rent payment deadline is December 15th. 3. The event in this call is: the landlord urging the tenant to pay the rent as soon as possible. 4. The result of this event is: the tenant promises to pay the rent to the landlord on December 15th.
[0127] 403. The terminal generates a summary of the call record based on the call record and the answer text of multiple preset questions through the summary generation model.
[0128] In this embodiment, the terminal processes the call log and the answer text of multiple preset questions using a summary generation model to generate a summary of the call log. Since the answer text of the multiple preset questions includes key information from the call log, generating a summary using the answer text and the call log not only ensures that the generated summary conforms to the original facts of the call log but also includes key information of interest to the user, thereby maximizing the accuracy and comprehensiveness of the summary.
[0129] In some embodiments, the terminal can also input preset task text into the summary generation model to instruct the model to generate a summary of the call log. The task text instructs the model to generate a summary of the call log based on the call log and the answer texts of multiple preset questions. The task text can be a prompt. Since the summary generation model already has the ability to generate fine-grained summaries that meet preset requirements, compared to the guidance text in the above embodiments, the task text does not need to describe the summary generation rules in detail, making its description simpler. For example, the task text could be: "Generate a summary of the call log based on the call log and the above text." The terminal can again invoke the summary generation model and input the call log, the answer texts of multiple preset questions, and the task text into the model to generate a summary of the call log according to the task text. By using simple task text to prompt the summary generation model about the task to be performed, the model's summary generation capabilities can be fully utilized to generate a fine-grained summary of the call log.
[0130] Because the terminal executes steps 401 and 402, it can complete the key information extraction and summary generation tasks in two separate stages using the summary generation model. Therefore, the summary generation model does not need to handle both tasks simultaneously, thus ensuring its performance on each task. Furthermore, by adopting this two-stage generation method, the key information extracted in the first stage can be referenced during the second stage of generating the call record summary, ensuring that the generated summary contains the key information from the call record and improving the controllability and accuracy of summary generation.
[0131] It should be noted that, in this embodiment, the processing of user call records is implemented through a model deployed on the client side. Therefore, generating call record summaries using a client-side summary generation model can effectively protect user privacy, reduce the risk of privacy leaks, and fully utilize the summary generation capabilities of the summary generation model to generate highly accurate call record summaries for users.
[0132] 404. Upon obtaining a summary of the call log, the terminal displays a second prompt message, which prompts whether to display the summary of the call log.
[0133] In this embodiment, after the terminal obtains a summary of the call log through the summary generation model, the terminal displays a second prompt message to indicate to the user whether to display the call log summary. It should be noted that steps 401-403 described above are imperceptible to the user of the terminal. After hanging up the call, the user can view the second prompt message displayed on the terminal. The user can choose to display or not display the call log summary according to their personal needs.
[0134] 405. In response to the confirmation operation of the second prompt message, the terminal displays a summary of the call log.
[0135] In this embodiment, in response to the user's confirmation of the second prompt information, the terminal displays a summary of the call log on the screen. The summary includes a summary of the call log and several core events within the call log. The summary may also include several pending events from the call log. Therefore, by displaying a summary of the call log to the user, multiple key pieces of information from the call log can be clearly and intuitively presented to the user, eliminating the need for the user to replay the call recording to obtain this key information, thus improving the efficiency of the user in obtaining key information.
[0136] This application provides a summary generation method. After a call ends, the terminal can extract multiple key pieces of information from the call log using a summary generation model. Leveraging the model's powerful fine-grained call log summary generation capabilities, the terminal processes the call log and key information to generate an accurate and comprehensive summary. The terminal then displays the call log summary to the user, allowing them to quickly grasp the core content of the call log by viewing the summary, thus improving the efficiency of reviewing key topics and tasks discussed in the call log.
[0137] All of the above-mentioned optional technical solutions can be combined in any way to form the optional embodiments of this application, and will not be described in detail here.
[0138] Figure 5 This is a schematic diagram of the structure of a training device for a summary generation model provided in an embodiment of this application. See also... Figure 5 The device includes: a first acquisition module 501, a first generation module 502, and a training module 503.
[0139] The first acquisition module 501 is used to acquire sample call records. The sample call records are marked with the call object and are the call content of the call object.
[0140] The first generation module 502 is used to process the sample call record through a natural language processing model based on the guiding text to obtain a first summary. The natural language processing model is used to process the input text data to obtain a summary of the text data. The guiding text is used to guide the natural language processing model to generate a summary of the sample call record in the form of a first-person subject. The summary contains multiple summary attributes, and the summary attributes are used to represent at least one key content in the sample call record.
[0141] Training module 503 is used to train the summary generation model based on sample call records and a first summary, so that the summary generation model can process call records to generate a summary of the call records.
[0142] In some embodiments, the training module 503 is used to process sample call records based on the summary generation model to obtain a second summary; determine the training loss of the summary generation model based on the first and second summaries; and train the summary generation model based on the training loss.
[0143] In some embodiments, Figure 6 This is a schematic diagram of the structure of a training device for another summary generation model provided in this application embodiment. See also... Figure 6 The device also includes:
[0144] The second generation module 504 is used to generate guiding text, which includes a first rule and a second rule. The first rule is used to define the multiple summary attributes that need to be included in the summary, and the second rule is used to define the generation rules and display format of the multiple summary attributes.
[0145] Input module 505 is used to input the guiding text and sample call records into the natural language processing model.
[0146] In some embodiments, the input module 505 is further configured to obtain a summary sample, wherein the summary sample is a call record summary that conforms to the first rule and the second rule; and input the summary sample, the guiding text, and the sample call record into the natural language processing model.
[0147] In some embodiments, the apparatus further includes:
[0148] Display module 506 is used to display a first prompt message when a first summary is obtained. The first prompt message is used to prompt whether the format and content of the first summary should be corrected.
[0149] The second acquisition module 507 is used to acquire the corrected first summary in response to the confirmation operation of the first prompt information.
[0150] In some embodiments, the apparatus further includes:
[0151] The pre-training module 508 is used to pre-train the summary generation model based on a pre-training dataset, which includes terms and nouns from multiple domains. The pre-training is used to enable the summary generation model to understand the terms and nouns from multiple domains in order to generate summaries of call records from multiple domains.
[0152] This application provides a training apparatus for a summary generation model. After obtaining sample call records with labeled call subjects, a natural language processing model processes the sample call records according to guiding text to obtain a summary with any call subject as the subject and containing multiple summary attributes. Then, the summary generation model is trained using the sample call records and their summaries, enabling the model to generate fine-grained call record summaries. That is, the summary generation model can generate summaries containing multiple summary attributes with any call subject in the call record as the subject, thereby ensuring that the summary contains the key content of the call record and improving the comprehensiveness and accuracy of the generated summary.
[0153] It should be noted that the training device for the summarization model provided in the above embodiments is only illustrated by the division of the above functional modules. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the terminal can be divided into different functional modules to complete all or part of the functions described above. In addition, the training device for the summarization model provided in the above embodiments and the training method embodiments for the summarization model belong to the same concept, and the specific implementation process can be found in the method embodiments, which will not be repeated here.
[0154] Figure 7 This is a schematic diagram of a summary generation apparatus provided in an embodiment of this application. See also... Figure 7 The device includes an input module 701, a first generation module 702, and a second generation module 703.
[0155] Input module 701 is used to input the call record and thought chain into the summary generation model in response to the end of the call. The call record is marked with the call object and is the call content of the call object. The thought chain is used to indicate multiple preset questions. The multiple preset questions are used to extract keywords and key sentences in the call record. The summary generation model is used to generate a summary of the call record in the form of a first-person subject. The summary contains multiple summary attributes, which are used to represent at least one key content in the call record.
[0156] The first generation module 702 is used to process the call record through a summary generation model based on the thought chain to obtain answer text for multiple preset questions. The answer text includes at least one of keywords and key sentences.
[0157] The second generation module 703 is used to generate a summary of the call record based on the call record and the answer text of multiple preset questions through a summary generation model.
[0158] In some embodiments, the input module 701 is configured to, in response to the end of a call, convert the call audio to text to obtain the call content; based on the caller and the called party, label the call content with the call objects to obtain the call record; and input the call record and thought chain into the summary generation model.
[0159] In some embodiments, the second generation module 703 is used to input call records, answer texts of multiple preset questions, and preset task text into a summary generation model, so as to generate a summary of the call records according to the task text through the summary generation model. The task text is used to instruct the summary generation model to generate a summary of the call records based on the call records and answer texts of multiple preset questions.
[0160] In some embodiments, Figure 8 This is a schematic diagram of another abstract generation device provided in an embodiment of this application. See also... Figure 8 The device also includes:
[0161] Display module 704 is used to display a second prompt message when a summary of the call log is obtained, the second prompt message being used to prompt whether to display the summary of the call log; and to display the summary of the call log in response to a confirmation operation of the second prompt message.
[0162] This application provides a summary generation device. After a call ends, the terminal can extract multiple key pieces of information from the call log using a summary generation model. Leveraging the model's powerful fine-grained call log summary generation capabilities, the terminal processes the call log and key information to generate an accurate and comprehensive summary. The terminal then displays the call log summary to the user, allowing them to quickly grasp the core content of the call log by viewing the summary, thus improving the efficiency of reviewing key topics and tasks discussed in the call log.
[0163] It should be noted that the summary generation device provided in the above embodiments is only illustrated by the division of the above functional modules. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the terminal can be divided into different functional modules to complete all or part of the functions described above. In addition, the summary generation device and the summary generation method embodiments provided in the above embodiments belong to the same concept, and the specific implementation process can be found in the method embodiments, which will not be repeated here.
[0164] This application provides a terminal, which includes a processor and a memory; the memory stores at least one piece of program code, which is executed by the processor to implement the training method of the summary generation model provided in the above method embodiments, or to implement the summary generation method provided in the above method embodiments.
[0165] Figure 9 This is a structural block diagram of a terminal provided in an embodiment of this application. In some embodiments, the terminal 900 is a smartphone, tablet computer, wearable device, or other terminal capable of accessing a wireless local area network as a wireless station. The terminal 900 in this application includes at least one or more of the following components: a processor 910, a memory 920, and at least two wireless links 930.
[0166] In some embodiments, the processor 910 includes one or more processing cores. The processor 910 connects to various parts within the terminal 900 using various interfaces and lines, and performs various functions and processes data of the terminal 900 by running or executing program code stored in the memory 920 and calling data stored in the memory 920. In some embodiments, the processor 910 is implemented using at least one hardware form of Digital Signal Processing (DSP), Field-Programmable Gate Array (FPGA), or Programmable Logic Array (PLA). The processor 910 can integrate one or more of the following: Central Processing Unit (CPU), Graphics Processing Unit (GPU), Neural-network Processing Unit (NPU), and modem. Specifically, the CPU primarily handles the operating system, user interface, and application programs; the GPU is responsible for rendering and drawing the content required for display on the screen; the NPU is used to implement Artificial Intelligence (AI) functions; and the modem is used for wireless communication. It is understandable that the aforementioned modem could also be implemented separately as a single chip without being integrated into the processor 910.
[0167] In some embodiments, the processor 910 is used to control the operating status of at least two wireless links 930. Accordingly, the processor 910 is a processor integrating a Wireless Fidelity (Wi-Fi) chip. This Wi-Fi chip is a chip with dual Wi-Fi processing capabilities. For example, the Wi-Fi chip is a dual-band dual-concurrent (DBDC) chip, or a dual-band simultaneous (DBS) chip, etc.
[0168] In some embodiments, the memory 920 includes random access memory (RAM), and in some embodiments, the memory 920 includes read-only memory (ROM). In some embodiments, the memory 920 includes non-transitory computer-readable storage medium. The memory 920 can be used to store program code. The memory 920 may include a program storage area and a data storage area, wherein the program storage area may store instructions for implementing an operating system, instructions for at least one function (such as touch function, sound playback function, image playback function, etc.), instructions for implementing the various method embodiments described below, etc.; the data storage area may store data created according to the use of the terminal 900 (such as audio data, phone book, etc.).
[0169] In some embodiments, the memory 920 stores reception schemes for different wireless links 930 receiving beacon frames, as well as identifiers of access nodes connected to different wireless links 930, identifiers of the wireless links 930, etc.
[0170] The at least two wireless links 930 are used to connect different access points (APs). They receive downlink data from the APs. These different access points can be access points within the same router or access points within different routers.
[0171] In some embodiments, the terminal 900 further includes a display screen. The display screen is a display component used to display a user interface. In some embodiments, the display screen is a touch-enabled display screen, allowing users to perform touch operations on the display screen using fingers, styluses, or any suitable object. In some embodiments, the display screen is typically located on the front panel of the terminal 900. In some embodiments, the display screen is designed as a full-screen, curved screen, irregularly shaped screen, dual-sided screen, or foldable screen. In some embodiments, the display screen is also designed as a combination of a full-screen and a curved screen, or a combination of an irregularly shaped screen and a curved screen, etc., which are not limited in this embodiment.
[0172] In addition, those skilled in the art will understand that the structure of the terminal 900 shown in the above figures does not constitute a limitation on the terminal 900. The terminal 900 may include more or fewer components than shown, or combine certain components, or have different component arrangements. For example, the terminal 900 may also include components such as a microphone, speaker, input unit, sensor, audio circuit, module, power supply, and Bluetooth module, which will not be described in detail here.
[0173] This application also provides a computer-readable storage medium storing at least one piece of program code, which is loaded and executed by the processor to implement the training method of the summary generation model shown in the above embodiments, or to implement the summary generation method shown in the above embodiments.
[0174] This application also provides a chip including programmable logic circuits and / or program instructions, which, when running on a terminal, are used to implement the training method of the summary generation model shown in the above embodiments, or to implement the summary generation method shown in the above embodiments.
[0175] This application also provides a computer program product that stores at least one piece of program code, which is executed by a processor to implement the training method of the summary generation model shown in the above embodiments, or to implement the summary generation method shown in the above embodiments.
[0176] The sequence numbers of the embodiments in this application are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.
[0177] Those skilled in the art will understand that all or part of the steps in the training method for implementing the summary generation model of the above embodiments can be implemented by hardware, or by a program instructing related hardware. The program can be stored in a computer-readable storage medium, such as a read-only memory, a disk, or an optical disk. The above descriptions are merely optional embodiments of this application and are not intended to limit this application. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of this application should be included within the protection scope of this application.
Claims
1. A training method for a summary generation model, characterized in that, The method includes: Obtain sample call records, wherein the sample call records are marked with call recipients and the sample call records are the call content of the call recipients; The sample call record is processed by a natural language processing model based on guiding text to obtain a first summary. The natural language processing model is used to process the input text data to obtain a summary of the text data. The guiding text is used to guide the natural language processing model to generate a summary of the sample call record in the form of a first-person subject. The summary contains multiple summary attributes, and the summary attributes are used to represent at least one key content in the sample call record. Based on the sample call records and the first summary, a summary generation model is trained so that the summary generation model can process the call records to generate a summary of the call records.
2. The method according to claim 1, characterized in that, The step of training the summary generation model based on the sample call records and the first summary includes: Based on the summary generation model, the sample call records are processed to obtain a second summary; Based on the first summary and the second summary, determine the training loss of the summary generation model; The summary generation model is trained based on the training loss.
3. The method according to claim 1, characterized in that, Before processing the sample call records based on the guiding text using a natural language processing model, the method further includes: Generate the guiding text, which includes a first rule and a second rule. The first rule is used to define multiple summary attributes that need to be included in the summary, and the second rule is used to define the generation rules and display format of the multiple summary attributes. The guiding text and the sample call record are input into the natural language processing model.
4. The method according to claim 3, characterized in that, The method further includes: Obtain a summary sample, wherein the summary sample is a call record summary that conforms to the first rule and the second rule; The summary example, the guiding text, and the sample call record are input into the natural language processing model.
5. The method according to claim 1, characterized in that, The method further includes: Upon obtaining the first summary, a first prompt message is displayed, which prompts whether the format and content of the first summary should be corrected. In response to the confirmation operation of the first prompt information, the corrected first summary is obtained.
6. The method according to claim 1, characterized in that, The method further includes: The summary generation model is pre-trained based on a pre-trained dataset, which includes terms and nouns from multiple domains. The pre-training is used to enable the summary generation model to understand the terms and nouns from the multiple domains in order to generate summaries of call records from the multiple domains.
7. A method for generating abstracts, characterized in that, The method includes: In response to the end of the call, the call log and thought chain of the call are input into the summary generation model. The call log is marked with the call subject and is the content of the call of the call subject. The thought chain is used to indicate multiple preset questions. The multiple preset questions are used to extract keywords and key sentences in the call log. The summary generation model is used to generate a summary of the call log in the form of a first-person subject. The summary contains multiple summary attributes. The summary attributes are used to represent at least one key content in the call log. The call record is processed using the summary generation model and based on the thought chain to obtain the answer text for the multiple preset questions. The answer text includes at least one of the keywords and key phrases. The summary generation model generates a summary of the call record based on the call record and the answer text of the multiple preset questions.
8. The method according to claim 7, characterized in that, The response to the end of the call includes inputting the call log and thought chain of the call into the summary generation model, including: In response to the end of the call, the audio of the call is converted to text to obtain the content of the call; Based on the caller and the called party in the call, the call content is labeled with the call subjects to obtain the call record; The call log and the thought chain are input into the summary generation model.
9. The method according to claim 7, characterized in that, The step of generating a summary of the call record using the summary generation model, based on the call record and the answer text of the multiple preset questions, includes: The call log, the answer texts of the multiple preset questions, and the preset task text are input into the summary generation model so that the summary generation model generates a summary of the call log according to the task text. The task text is used to instruct the summary generation model to generate a summary of the call log based on the call log and the answer texts of the multiple preset questions.
10. The method according to claim 7, characterized in that, The method further includes: Upon obtaining a summary of the call log, a second prompt message is displayed, which prompts whether to display the summary of the call log. In response to the confirmation of the second prompt message, a summary of the call record is displayed.
11. A training device for a summary generation model, characterized in that, The device includes: The first acquisition module is used to acquire sample call records, wherein the sample call records are marked with call objects and the sample call records are the call content of the call objects; The first generation module is used to process the sample call record using a natural language processing model based on guiding text to obtain a first summary. The natural language processing model is used to process the input text data to obtain a summary of the text data. The guiding text is used to guide the natural language processing model to generate a summary of the sample call record in the form of a first-person subject. The summary contains multiple summary attributes, and the summary attributes are used to represent at least one key content in the sample call record. The training module is used to train the summary generation model based on the sample call records and the first summary, so that the summary generation model can process the call records to generate a summary of the call records.
12. A summary generation apparatus, characterized in that, The device includes: An input module is used to input the call log and thought chain of the call into a summary generation model in response to the end of the call. The call log is marked with the call subject and is the content of the call of the call subject. The thought chain is used to indicate multiple preset questions. The multiple preset questions are used to extract keywords and key sentences from the call log. The summary generation model is used to generate a summary of the call log in the form of a first-person subject. The summary contains multiple summary attributes, and the summary attributes are used to represent at least one key content in the call log. The first generation module is used to process the call record based on the thought chain through the summary generation model to obtain the answer text of the multiple preset questions, wherein the answer text includes at least one of the keywords and the key sentences; The second generation module is used to generate a summary of the call record based on the call record and the answer text of the multiple preset questions through the summary generation model.
13. A computer device, characterized in that, The computer device includes a processor and a memory; the memory stores at least one piece of program code, which is executed by the processor to implement the training method of the summarization generation model as described in any one of claims 1 to 6, or to implement the summarization generation method as described in any one of claims 7 to 10.
14. A computer-readable storage medium, characterized in that, The storage medium stores at least one piece of program code, which is executed by a processor to implement the training method of the summarization generation model as described in any one of claims 1 to 6, or to implement the summarization generation method as described in any one of claims 7 to 10.
15. A computer program product, comprising a computer program, characterized in that, The computer program product stores at least one piece of program code, which is executed by a processor to implement the training method of the summarization generation model as described in any one of claims 1 to 6, or to implement the summarization generation method as described in any one of claims 7 to 10.
Citation Information
Patent Citations
Model training method and device, electronic equipment and readable storage medium
CN116991975A
Natural language text generation using semantic objects
US20200356732A1