Self-supervised abstract generation method, system, equipment and medium

Through the self-supervised abstract generation method, the teacher model is optimized using the self-supervised filter and expert iterative module, and the knowledge is distilled into the student model, solving the problem of low quality of the existing automatic abstract method to generate abstracts, achieving high-quality, controllable and efficient abstract generation.

CN120196750AActive Publication Date: 2025-06-24SHANG HAI JIE YUE XING CHEN ZHI NENG KE JI YOU XIAN GONG SI
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510677254.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-26
Publication Date
2025-06-24
Estimated Expiration
2045-05-26

AI Technical Summary

Technical Problem

The digests generated by existing automatic digest methods are of low quality, lack unified quantitative standards, and relying on large-scale pre-trained language models leads to high computing resources.

Method used

The self-supervised abstract generation method is adopted to generate documents-abstract pairs through pre-training language models, and the data set is filtered using the self-supervised filter module. The expert iterative module performs multiple rounds of self-supervised learning, iteratively optimizes the teacher model, and distillates the knowledge into the student model. The control attribute is introduced to generate customized abstracts through the controllable abstract generation module.

Benefits of technology

Reliance on large-scale pre-trained models is reduced, controllability, quality and cross-domain adaptability of the summary, and significantly reduce the computational overhead of training and inference.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120196750A_ABST
    Figure CN120196750A_ABST
Patent Text Reader

Abstract

The embodiment of the invention relates to the technical field of natural language processing, and discloses a self-supervised abstract generation method and system, equipment and a medium. Using a pre-training language model to generate a document-abstract pair; screening the document-abstract pair through a self-supervised filter module to obtain a data set; performing multi-round self-supervised learning on the data set by using an expert iteration module, and iterating a teacher model; distilling the iterated knowledge of the teacher model into a student model; and introducing control attributes in training and reasoning stages through a controllable abstract generation module, and generating a customized abstract by using the student model according to the control attributes. A self-supervised information theory target is constructed, and distillation and training are performed in combination with a small language model, so that the dependence on a large-scale pre-training model is reduced. The technical problem of low abstract generation quality can be at least solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of natural language processing, and in particular, to a self-supervised abstract generation method, system, device, and medium. Background Art

[0002] Automatic summarization technology is an important task in natural language processing (NLP), aiming to extract concise and informative content from the original text. Existing automatic summarization methods can be classified into the following categories: Extractive Summarization: This method selects some important sentences or paragraphs from the original text as the summary. Common techniques include those based on TF-IDF, TextRank algorithm, etc. The advantage of this method is that the generated summary is linguistically more consistent and easy to understand. However, its limitation is that it cannot perform effective content compression, and the generated summary may lack fluency and naturalness.

[0003] Generative Summarization: The generative method aims to understand the content of the original text and restate it in a more concise language. This method relies on deep learning models, especially sequence-to-sequence models (Seq2Seq) and pre-trained language models (such as BERT, GPT, etc.). Generative summary models have strong content compression capabilities and flexibility, and can generate more fluent and natural text. However, such methods usually require a large amount of manually annotated data for training, and the training process is complex and consumes a large amount of computing resources.

[0004] In practical applications, most generative summarization methods rely on large-scale pre-trained language models such as GPT-4 and ChatGPT, and rely on unsupervised or imitation learning for training. However, there is a lack of a unified quantitative standard in the prior art for what kind of summary is considered "good". Summary of the Invention

[0005] An object of this application is to provide a self-supervised abstract generation method, device, medium, and product, at least to solve the technical problem of low-quality abstract generation.

[0006] To achieve the above object, some embodiments of this application provide the following aspects: In a first aspect, some embodiments of this application further provide a self-supervised abstract generation method, including: using a pre-trained language model to generate document-abstract pairs; screening the document-abstract pairs through a self-supervised filter module to obtain a data set; using an expert iteration module to perform multiple rounds of self-supervised learning on the data set to iterate the teacher model; distilling the knowledge of the iterated teacher model into a student model; and introducing a control attribute through a controllable abstract generation module in the training and inference stages, and using the student model to generate a customized abstract according to the control attribute.

[0007] In a second aspect, some embodiments of the present application further provide a self-supervised abstract generation system, including: a pre-trained language model for generating document-abstract pairs; a self-supervised filter module for screening the document-abstract pairs to obtain a data set; an expert iteration module for performing multiple rounds of self-supervised learning on the data set to iteratively train a teacher model; distilling the knowledge of the iterated teacher model into a student model; and a controllable abstract generation module for introducing control attributes in the training and inference phases and using the student model to generate customized abstracts according to the control attributes.

[0008] In a third aspect, some embodiments of the present application further provide an electronic device, which includes: one or more processors; and a memory storing computer program instructions, where the computer program instructions, when executed, cause the processors to execute the steps of the method described above.

[0009] In a fourth aspect, some embodiments of the present application further provide a computer-readable medium having computer program instructions stored thereon, where the computer program instructions can be executed by a processor to implement the method described above.

[0010] Compared with the related art, in the solution provided by the embodiments of the present application, by constructing a self-supervised information-theoretic objective and combining distillation and training with a small language model, not only the dependence on large-scale pre-trained models is reduced, but also the quality of the model is optimized through information maximization and expert iteration. In addition, the present application significantly improves the controllability, quality, and cross-domain adaptability of the abstract by introducing control attributes and a self-supervised filtering mechanism. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] One or more embodiments are illustrated by way of example in the accompanying drawings, which do not constitute a limitation on the embodiments. Elements with the same reference numerals in the drawings are represented as similar elements, unless otherwise stated, and the drawings in the figures do not constitute a scale limitation.

[0012] Figure 1 It is a schematic flowchart of a self-supervised abstract generation method provided according to an embodiment of the present application; Figure 2 It is a schematic structural diagram of a self-supervised abstract generation system provided according to an embodiment of the present application; Figure 3 It is a schematic structural diagram of an electronic device provided according to an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0013] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the following will clearly and completely describe the technical solutions in the embodiments of this application with reference to the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are some, but not all, of the embodiments of this application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of this application without creative efforts fall within the scope of protection of this application.

[0014] First Embodiment The first embodiment of this application relates to a self-supervised abstract generation method. As Figure 1 shown, the method may include the following steps: S101, using a pre-trained language model to generate document-abstract pairs: Provide a basic pre-trained language model. Although it does not directly have the specialized abstract generation ability itself, it can generate preliminary documents and their corresponding abstracts based on its powerful language understanding and generation capabilities. This model is used as a teacher model to generate preliminary document-abstract pairs.

[0015] S102, screening the document-abstract pairs through a self-supervised filter module to obtain a dataset: Use a self-supervised language model (such as the BERT model based on the masked language model MLM) to calculate the mutual information (PMI) between the document and the abstract. PMI measures the degree of association between the document and the abstract. Set the screening criteria from three aspects: significance, credibility, and conciseness: Significance: Ensure that the abstract contains the key information in the document, and measure the significance by calculating the overlap degree between the keywords in the abstract and the keywords in the document. For example, set a threshold for the keyword overlap ratio. Only when the keyword overlap ratio between the abstract and the document exceeds this threshold is the abstract considered to have sufficient significance.

[0016] Credibility: Ensure that the abstract content is true and reliable and does not contain fictional information. The credibility can be evaluated by checking whether the factual information in the abstract is consistent with the document. For example, for an abstract containing specific data and events, verify whether there is a corresponding accurate description in the document.

[0017] Conciseness: Control the length of the abstract to avoid containing too much redundant information. A ratio threshold of the abstract length to the document length can be set to screen out the abstracts that meet the length requirements.

[0018] Screen the document-abstract pairs according to the calculated PMI value and the set screening criteria, and screen out the document-abstract pairs with a PMI value higher than the preset threshold and meeting the requirements of significance, credibility, and conciseness to form a high-quality dataset.

[0019] S103. Use the expert iteration module to perform multiple rounds of self-supervised learning on the dataset to iterate the teacher model: Use the expert iteration module to perform multiple rounds of self-supervised learning on the dataset to iteratively optimize the teacher model. The teacher model can select a model with a Transformer architecture, such as the T5 model. In each round of self-supervised learning, input the document-summary pairs in the dataset into the teacher model for forward propagation calculation. The teacher model will generate a predicted summary based on the input document, and then compare the predicted summary with the true summary in the dataset to calculate the loss function between the two.

[0020] A commonly used loss function can select the cross-entropy loss function. According to the calculated loss function, use the backpropagation algorithm to update the parameters of the teacher model. Through multiple rounds of self-supervised learning and parameter updates, gradually improve the summary generation ability of the teacher model.

[0021] S104. Distill the knowledge of the iterated teacher model into the student model: To reduce the computational cost and improve the inference efficiency of the model, use information distillation technology to distill the knowledge of the iterated teacher model into a small student model. The student model can select a Transformer model with a smaller parameter scale, such as DistilBERT. Specifically, define the loss function between the outputs of the teacher model and the student model, and this loss function measures the difference between their outputs. A commonly used loss function can be the mean squared error loss function. By minimizing this loss function, adjust the parameters of the student model so that the output of the student model is as close as possible to the output of the teacher model, thereby transferring the knowledge of the teacher model to the student model.

[0022] S105. Through the controllable summary generation module, introduce control attributes in the training and inference stages, and use the student model to generate customized summaries according to the control attributes: Use the controllable summary generation module to introduce control attributes in the training and inference stages of the model to generate customized summaries that meet the specific needs of users. The control attributes can include the length of the summary, the extractiveness of the summary, the keywords that the summary focuses on, etc.

[0023] In the training stage, input the control attributes together with the document-summary pairs into the student model. For example, for the control attribute of the summary length, the desired summary length can be encoded as a vector and input into the student model together with the document. The model will learn the ability to generate appropriate summaries under different control attributes.

[0024] In the inference stage, according to the control attributes input by the user, the trained student model is used to generate a customized abstract that meets the requirements. For example, if the user hopes to generate an abstract with a length of 200 words and focuses on a specific keyword in the document, these control attributes are input into the student model, and the student model will generate an abstract that meets the requirements based on this information.

[0025] It is not difficult to find that, compared with the related technologies, in the solution provided by the embodiments of the present application, through the information distillation technology, the knowledge of the optimized teacher model is transferred to the small student model, reducing the dependence on the large pre-trained model. While maintaining the high quality of the abstract, the computational overhead of training and inference is greatly reduced, effectively solving the problem of high computational cost of traditional methods; with the help of the information maximization objective, the abstract generation is optimized from three aspects of saliency, credibility and conciseness, and the teacher model is subjected to multiple rounds of self-supervised learning through the expert iteration mechanism, gradually improving the model generation ability, so that the generated abstract can accurately extract the key information of the document, faithfully reflect the original content and be concise; by introducing control attributes, users can flexibly specify personalized requirements such as abstract length, extractability, keyword focus, etc. The model learns the generation rules under different control attributes during the training stage and can dynamically adjust the generation strategy according to the user input during inference, significantly enhancing the controllability and flexibility of abstract generation.

[0026] Second Embodiment The second embodiment of the present application relates to a self-supervised abstract generation method. The second implementation is an improvement based on the first embodiment. The specific improvement lies in: The screening of the document-abstract pair by the self-supervised filter module includes: performing a mask process on the document-abstract pair, calculating the mutual information between the document and the abstract, and calculating the saliency score, credibility score and conciseness score of the abstract according to the mutual information; screening out the document-abstract pairs whose saliency score, credibility score and conciseness score meet the preset threshold values to form the data set.

[0027] The self-supervised filter module is used to screen and optimize the preliminarily generated document-abstract pairs to ensure that they meet the requirements of saliency, credibility and conciseness. Use a self-supervised language model (MLM) to perform a mask process on the document, calculate the mutual information (PMI) between the document and the abstract, and evaluate the saliency, credibility and conciseness of the abstract.

[0028] Saliency: Calculate the correlation between the document and the abstract through the mutual information formula (PMI) to ensure that the abstract contains the key information in the document. Saliency score formula: ;

[0029] x represents the document; y represents the abstract; [] is an indicator function (1 when the condition holds, 0 otherwise); is the probability that the masked language model (MLM) predicts the masked part as the original content of x when part of the content of document x is masked ( ) and the known abstract y; is the probability that the MLM predicts the masked part as the original content of x when only the masked part of the document is known (without the information of abstract y).

[0030] The formula determines the prominence of the abstract for the key information of the document by comparing the absolute value of the logarithmic ratio of the predicted probabilities of the masked part of the document with and without the abstract y. If this value is greater than the threshold , then indicates that the abstract significantly contains the key information of the document; otherwise it is 0.

[0031] Faithfulness: Ensure that no information not in the original text is added to the abstract by the possibility of restoring the abstract content. Faithfulness scoring formula: ; is the probability that the MLM predicts the masked part as the original content of y when part of the content of abstract y is masked and the known document x; is the probability that the MLM predicts the masked part as the original content of y when only the masked part of the abstract is known (without the information of document x).

[0032] This formula measures the degree of support of the document for the content of the abstract. If the absolute value of the logarithmic ratio of the predicted probabilities of the masked part of the abstract with and without the document x is greater than the threshold , then indicates that the content of the abstract is reliable based on the document; otherwise it is 0.

[0033] Brevity: Ensure that the abstract is concise and refined according to the length control of the abstract. Brevity scoring formula: ; is the length of the abstract (such as the number of words, characters), is the length of the document.

[0034] The formula determines the brevity by the ratio of the length of the abstract to the length of the document. If ( is the preset length ratio threshold), then It indicates that the abstract is concise and has no excessive redundant information; otherwise, it is 0.

[0035] Only retain the document-abstract pairs that meet the requirements in terms of significance, credibility, and conciseness. After screening, a high-quality dataset is obtained. Generally speaking, the core point is to calculate the PMI between the abstract y and the original text x (i.e., the probability of mutual information between the two), and whether the ratio of the scores between the two reaches a certain threshold. If it reaches, it means the abstract is useful, and optimization has been carried out from three aspects.

[0036] Further, the iterative teacher model includes: inputting the document-abstract pairs in the dataset into the teacher model for forward propagation calculation; calculating the first loss function according to the difference between the output result of the teacher model and the abstract; and updating the parameters of the teacher model using the backpropagation algorithm according to the first loss function.

[0037] Optimize the teacher model through multiple rounds of iteration to generate higher-quality abstract data. Based on self-supervised learning, use the self-generated high-quality abstract pairs to train the teacher model to improve its generation ability in each round of iteration.

[0038] Optimization process: Assume that the current teacher model generates preliminary document-abstract pairs (x, y), and optimize the teacher model through the following steps: Data generation: ; Data screening: Screen out high-quality document-abstract pairs through a self-supervised filter to obtain a dataset; Model optimization: Update the teacher model using the screened dataset:

[0039] Represents the state of the teacher model after the (t + 1)-th round of iteration; Represents finding the teacher model that maximizes the subsequent expectation ; E(x, y) takes the expectation of the samples (x, y) in the high-quality dataset ; is the teacher model The logarithm of the probability of generating the abstract y given the document x. In each round of iteration, use the high-quality dataset to update the parameters of the teacher model by maximizing "the expectation of the logarithm of the probability that the model generates the corresponding abstract y given the document x". After multiple rounds of iteration, the teacher model can better learn the ability to generate high-quality abstracts according to the document, improving the accuracy and quality of abstract generation.

[0040] Further, the distillation of the knowledge of the iterated teacher model into the student model includes: defining a second loss function between the outputs of the teacher model and the student model, and calculating the difference value between the two outputs according to the second loss function; adjusting the parameters of the student model by minimizing the difference value of the loss function output.

[0041] Transfer the knowledge of the teacher model optimized by expert iteration to a smaller student model to achieve efficient inference. Through information distillation technology, the high-quality document-summary pairs generated by the teacher model are used as training data to optimize the student model. The distillation process: Generate data:

[0042] Train the student model: Optimize the student model through the following objective function:

[0043] represents the student model, which is a lightweight model designed to achieve the abstract generation function at a low computational cost; represents in the parameter space of the student model to find the model parameter configuration that maximizes the subsequent expectation; E(x, ) takes the expectation of the sample set, reflecting the comprehensive consideration of multiple groups of input-output samples; is the student model When given the document x, generate the abstract generated by the teacher model. By maximizing this value, the student model learns to imitate the abstract generation behavior of the teacher model.

[0044] Using knowledge distillation technology, let the student model learn the abstract generation ability of the teacher model. The generated by the teacher model is used as a supervision signal. The student model optimizes its own parameters to maximize the expected value of the probability logarithm of generating , so as to transfer the knowledge of the teacher model to itself, and improve the abstract generation performance while maintaining a low computational cost. The student model can generate high-quality abstracts similar to the teacher model with a lower computational overhead.

[0045] Further, the control attributes include: abstract length, information extraction, and keywords.

[0046] Generate abstracts in a specific style according to user needs, such as controlling the length, extractability, or keyword focus of the abstract. Control attribute tags: Add control attributes such as the length, information extraction, keywords, etc. of the abstract to the document-summary pair during the training phase. Pass these control attributes as input conditions to the model to ensure that the generated abstract meets the predetermined requirements.

[0047] Application examples of control attributes: abstract length control, control_attributes = {length: medium}; information extraction control, control_attributes = {extractiveness: high}; keyword focus control, control_attributes = {focus_keyword: disease}. The generated abstract y will meet these control attribute requirements to ensure that the abstract meets the user's needs.

[0048] Furthermore, using the student model to generate a customized abstract according to the control attributes includes: obtaining a user input instruction, extracting the input instruction to obtain the control attributes; inputting the control attributes into a pre-trained language model, and passing them to the self-supervised filter module, the expert iteration module, and the controllable abstract generation module to generate the customized abstract.

[0049] The user will input an instruction containing specific requirements for generating the abstract, such as "generate an abstract about the development of artificial intelligence that is about 200 words long and highlights the key technological innovation points". The system will perform natural language processing on the input instruction to extract the key control attributes. In the above example, the control attributes include "abstract length is about 200 words" and "highlight the key technological innovation points".

[0050] Input the extracted control attributes into the pre-trained language model. The pre-trained language model will encode the control attributes and convert them into a vector representation suitable for subsequent module processing. The self-supervised filter module will perform more accurate screening on the generated document-abstract pairs according to the control attributes. For example, if the control attribute requires the abstract length, this module will screen out the document-abstract pairs that meet the length requirement.

[0051] When training the teacher model, the expert iteration module will optimize the model in combination with the control attributes. For example, if the control attribute emphasizes certain keywords, the model will pay more attention to the content related to these keywords. The controllable abstract generation module directly uses the control attributes to generate a customized abstract. It will adjust the strategy for generating the abstract according to the control attributes to ensure that the generated abstract meets the specific needs of the user.

[0052] Furthermore, using the student model to generate a customized abstract according to the control attributes includes: In the training stage, according to the control attributes and the document-abstract pairs, guide the student model to learn to generate corresponding abstracts under the control attributes; in the inference stage, according to the control attributes, use the trained student model to generate the customized abstract.

[0053] During the training phase, the system combines control attributes with document-summary pairs. For example, for each document-summary pair, the corresponding control attributes (such as summary length, keywords, etc.) are input simultaneously. The student model will learn the ability to generate corresponding summaries under different control attributes. Through a large number of training samples, the model will gradually master how to adjust the summary generation method according to the control attributes to meet various requirements.

[0054] In the inference phase, summaries are generated according to the control attributes. When receiving new documents and control attributes, the trained student model is used to generate customized summaries. The student model will generate compliant summaries according to the input control attributes by applying the knowledge learned during the training phase.

[0055] It is not difficult to find that in the embodiments of the present application, by extracting control attributes from user input and utilizing these attributes in the training and inference phases, the system can generate customized summaries that meet the diverse needs of users. This method enhances the flexibility and practicality of summary generation, enabling the system to better adapt to applications in different scenarios.

[0056] The step division of the above various methods is only for clear description. When implemented, they can be combined into one step or some steps can be split into multiple steps. As long as the same logical relationship is included, they are all within the protection scope of the present application; adding insignificant modifications to the algorithm or process or introducing insignificant designs, but without changing the core design of the algorithm and process, are all within the protection scope of this application.

[0057] The Third Embodiment The third embodiment of the present application relates to a self-supervised summary generation system / device. As Figure 2 shown, the system includes: A pre-trained language model for generating document-summary pairs; A self-supervised filter module for screening the document-summary pairs to obtain a data set; An expert iteration module for performing multiple rounds of self-supervised learning on the data set to iterate the teacher model; and distilling the knowledge of the iterated teacher model into the student model; A controllable summary generation module for introducing control attributes in the training and inference phases and using the student model to generate customized summaries according to the control attributes.

[0058] It is not difficult to find that this embodiment is a system embodiment corresponding to the first embodiment, and this embodiment can be implemented in cooperation with the first embodiment. The relevant technical details mentioned in the first embodiment are still valid in this embodiment. To avoid repetition, they are not elaborated here. Correspondingly, the relevant technical details mentioned in this embodiment can also be applied in the first embodiment.

[0059] It is worth mentioning that each module involved in this embodiment is a logical module. In practical applications, a logical unit can be a physical unit, a part of a physical unit, or implemented as a combination of multiple physical units. In addition, to highlight the innovative part of this application, units that are not closely related to solving the technical problems proposed in this application are not introduced in this embodiment, but this does not mean that there are no other units in this embodiment.

[0060] In addition, some embodiments of this application also provide an electronic device. The electronic device can be various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and so on. The electronic device can also be various forms of mobile devices, such as personal digital processors, cellular phones, smart phones, wearable devices, and other similar computing devices.

[0061] The electronic device includes: one or more processors; and a memory storing computer program instructions that, when executed, cause the processors to perform the steps of the method provided in any one or more of the above embodiments. Figure 3 An exemplary structural diagram of the electronic device is disclosed. As Figure 3 shown, the electronic device includes: one or more processors 1101, a memory 1102, and interfaces for connecting the components, including a high-speed interface and a low-speed interface. Each component is interconnected using different buses and can be installed on a common motherboard or in other ways as needed. The processor can process instructions executed within the electronic device, including instructions stored in the memory or on the memory to display graphical information of the GUI on an external input / output device (such as a display device coupled to the interface). In some other embodiments, if necessary, multiple processors and / or multiple buses can be used together with multiple memories and multiple memories. Similarly, multiple electronic devices can be connected, and each device provides some necessary operations (for example, as a server array, a set of blade servers, or a multi-processor system). Among them, the components, their connections and relationships, and their functions shown herein are only examples and are not intended to limit the implementation of this application described and / or claimed herein.

[0062] The electronic device may further include: an input device 1103 and an output device 1104. The processor 1101, the memory 1102, the input device 1103, and the output device 1104 can be connected by a bus or other means, Figure 3 taking connection by bus as an example.

[0063] The input device 1103 can receive input digital or character information and generate key signal inputs related to the user settings and function controls of the electronic device, such as input devices like touchscreens, keypads, mice, trackpads, touchpads, pointing sticks, one or more mouse buttons, trackballs, joysticks, etc. The output device 1104 can include display devices, auxiliary lighting devices (e.g., LEDs), and tactile feedback devices (e.g., vibration motors), etc. The display device can include, but is not limited to, liquid crystal displays (LCDs), light-emitting diode (LED) displays, and plasma displays. In some embodiments, the display device can be a touchscreen.

[0064] To provide interaction with the user, the electronic device can be a computer. The computer has: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and a pointing device (e.g., a mouse or a trackball), through which the user can provide inputs to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and inputs from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0065] In the embodiments of the present application, a computer program / instructions is stored on a computer-readable medium. When the computer program / instructions are executed by a processor, the steps of the method provided in any one or more of the above embodiments are implemented. The computer-readable medium can be included in the electronic device described in the above embodiments; or it can exist separately without being assembled into the device. The above computer-readable medium carries one or more computer-readable instructions.

[0066] The memory 1102 can be used as a non-transitory computer-readable storage medium and can be used to store non-transitory software programs, non-transitory computer-executable programs, and modules. The processor 1101 executes various functional applications and data processing of the server by running the non-transitory software programs, instructions, and modules stored in the memory 1102, so as to implement the program instructions / modules corresponding to the method provided in any one or more of the above embodiments of the present application.

[0067] The memory 1102 may include a program storage area and a data storage area. The program storage area may store an operating system and application programs required for at least one function. The data storage area may store data created according to the use of the electronic device and the like. In addition, the memory 1102 may include a high-speed random access memory, and may also include a non-transitory memory, such as at least one magnetic disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some embodiments, the memory 1102 may optionally include a memory remotely disposed relative to the processor 1101, and these remote memories may be connected to the electronic device through a network. Examples of the above network include but are not limited to the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0068] It should be noted that the computer-readable medium described in this application may be a computer-readable signal medium, a computer-readable storage medium, or any combination of the above two. The computer-readable medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium may include but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In this application, the computer-readable medium may be any tangible medium that contains or stores a program, and this program may be used by or in combination with an instruction execution system, apparatus, or device.

[0069] The computer-readable medium includes permanent and non-permanent, removable and non-removable media, and information storage can be implemented by any method or technology. The information may be computer-readable instructions, data structures, program modules, or other data. Examples of the computer's storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory, or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD), or other optical storage, magnetic cassette tapes, magnetic disk storage, or other magnetic storage devices, or any other non-transmission medium that can be used to store information accessible by a computing device.

[0070] Computer program code for performing the operations of this application can be written in one or more programming languages or combinations thereof. The programming languages include object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computer (e.g., connected through the Internet using an Internet service provider).

[0071] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. For example, an application-specific integrated circuit (ASIC), a general-purpose computer, or any other similar hardware device can be used. In some embodiments, the software program of this application can be executed by a processor to implement the above steps or functions. Similarly, the software program of this application (including related data structures) can be stored in a computer-readable recording medium, such as a RAM memory, a magnetic or optical drive, or a floppy disk and similar devices. Additionally, some steps or functions of this application can be implemented using hardware, for example, as a circuit that cooperates with a processor to execute each step or function.

[0072] The computer program product provided by the embodiments of this application includes one or more computer programs / instructions. When the computer programs / instructions are executed by a processor, they wholly or partly generate the processes or functions described in the embodiments of this application. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center by wire (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wirelessly (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server, data center, etc. that contains one or more integrated available media. The available media can be magnetic media (e.g., floppy disk, hard disk, magnetic tape), optical media (e.g., DVD), or semiconductor media (e.g., solid state disk (SSD)), etc.

[0073] The flowcharts or block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of devices, methods, and computer program products according to various embodiments of the present application. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a part of code, which contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than that marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and combinations of blocks in the block diagram and / or flowchart, may be implemented by a dedicated hardware-based system that performs the specified functions or operations, or may be implemented by a combination of dedicated hardware and computer instructions.

[0074] The scope of the present application is defined by the appended claims rather than the above description. Therefore, all changes that fall within the meaning and scope of the equivalent elements of the claims are intended to be encompassed within the present application. Any reference signs in the claims should not be construed as limiting the claims involved. In addition, it is obvious that the word "comprising" does not exclude other units or steps, and the singular does not exclude the plural. The multiple units or devices stated in the apparatus claims may also be implemented by one unit or device through software or hardware. The words "first", "second", etc. are only used for distinguishing descriptions and do not represent any specific order, nor can they be construed as indicating or implying relative importance.

[0075] As described above, the above are only specific embodiments of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present application can easily mention changes or substitutions, which should all be covered within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims, and the above embodiments should be regarded as exemplary and non-limiting.

Claims

1. A self-supervised abstract generation method, characterized in that, The method includes: Using a pre-trained language model to generate document-summary pairs; Filtering the document-summary pairs through a self-supervised filter module to obtain a dataset; Using an expert iteration module to perform multiple rounds of self-supervised learning on the dataset and iterate the teacher model; Distilling the knowledge of the iterated teacher model into a student model; Through a controllable summary generation module, introducing control attributes in the training and inference phases, and using the student model to generate customized summaries according to the control attributes.

2. The method according to claim 1, wherein The filtering of the document-summary pairs through the self-supervised filter module includes: Masking the document-summary pairs, calculating the mutual information between the document and the summary, and calculating the significance score, credibility score, and conciseness score of the summary according to the mutual information; Filtering out the document-summary pairs whose significance score, credibility score, and conciseness score meet the preset thresholds to form the dataset.

3. The method according to claim 2, wherein The iterated teacher model includes: Inputting the document-summary pairs in the dataset into the teacher model for forward propagation calculation; Calculating a first loss function according to the difference between the output result of the teacher model and the summary; Updating the parameters of the teacher model using the backpropagation algorithm according to the first loss function.

4. The method according to claim 3, wherein The distilling of the knowledge of the iterated teacher model into the student model includes: Defining a second loss function between the outputs of the teacher model and the student model, and calculating the difference value between their outputs according to the second loss function; Adjusting the parameters of the student model by minimizing the difference value of the output of the loss function.

5. The method according to any one of claims 1 to 4, characterized in that, The control attributes include: summary length, information extraction, and keywords.

6. The method according to claim 5, characterized in that, The using the student model to generate customized summaries according to the control attributes includes: Obtaining a user input instruction, extracting the input instruction to obtain the control attributes; Inputting the control attributes into the pre-trained language model, passing them to the self-supervised filter module, the expert iteration module, and the controllable summary generation module to generate the customized summaries.

7. The method according to claim 6, characterized in that, The using the student model to generate customized summaries according to the control attributes includes: In the training phase, guiding the student model to learn to generate corresponding summaries under the control attributes according to the control attributes and the document-summary pairs; In the inference phase, using the trained student model to generate the customized summaries according to the control attributes.

8. A self-supervised abstract generation system, characterized in that, The system includes: A pre-trained language model for generating document-summary pairs; A self-supervised filter module for filtering the document-summary pairs to obtain a dataset; An expert iteration module for performing multiple rounds of self-supervised learning on the dataset, iterating the teacher model; and distilling the knowledge of the iterated teacher model into the student model; A controllable summary generation module for introducing control attributes in the training and inference phases and using the student model to generate customized summaries according to the control attributes.

9. An electronic device, characterized in that, The electronic device includes: One or more processors; and A memory storing computer program instructions, which when executed cause the processor to execute the steps of the method according to any one of claims 1 to 7.

10. A computer-readable medium having computer programs / instructions stored thereon, characterized in that, When the computer program / instructions are executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.

Citation Information

Patent Citations

  • Knowledge distillation-based text classification method and system

    CN114818902A

  • Abstract information generation and search result display method and device, equipment and medium

    CN115080816A

  • Lightweight abstract generation method based on knowledge extraction

    CN116341541A

  • Code abstract generation method, system and equipment based on improved Transform model

    CN118963825A

  • Autonomous and continuously self-improving learning system

    US11100373B1