Text generation method and device, computer equipment and computer readable storage medium
By physically processing the text information to be processed and generating entity prompt information, the problem of poor summary quality in the existing generative summary task is solved, and the generated summary text information is achieved with higher quality and more relevant content.
Patent Information
- Application Number
- CN202311406176.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-10-26
- Publication Date
- 2025-05-02
AI Technical Summary
The digests generated in existing generative summary tasks are of poor quality, with problems that are not related to long text or have duplicate content.
By performing entity processing on the text information to be processed, the first entity prompt information is obtained, and the target text information is generated based on the text information to be processed and the first entity prompt information is used to reduce the situation where the generated content does not match or is duplicated with the original text.
The quality of generated summary text information is improved, and the situations inconsistent with the content of the pending text information or duplication of the content is reduced.
Smart Images

Figure CN119917658A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of natural language processing, and in particular to a text generation method, apparatus, computer device and computer-readable storage medium. Background Art
[0002] The generative summarization task refers to the task of generating a short summary based on a long text. In the era of news information explosion, the generative summarization task has attracted widespread attention because it can provide readers with convenience to quickly select the news information they need. However, the existing generative summarization tasks have the problem of poor quality of generated summaries. Summary of the invention
[0003] The embodiments of the present application provide a text generation method, apparatus, computer device, and computer-readable storage medium, which can improve the quality of generated summary text information.
[0004] The technical solution adopted by the present invention to solve the problem is as follows:
[0005] On the one hand, the present application provides a text generation method, comprising:
[0006] Get the text information to be processed;
[0007] Perform entity processing on the text information to be processed to obtain first entity prompt information;
[0008] Target text information is generated based on the text information to be processed and the first entity prompt information.
[0009] In some implementation schemes of the present application, entity processing is performed on the text information to be processed to obtain first entity prompt information, including:
[0010] Perform entity extraction on the text information to be processed to obtain multiple candidate entity information of the text information to be processed;
[0011] Determine first entity information of each candidate entity information; the first entity information represents the importance of each candidate entity information;
[0012] First entity prompt information is determined based on the first entity information and the plurality of candidate entity information.
[0013] In some embodiments of the present application, determining the first entity information of each candidate entity information includes:
[0014] Obtaining first position information of each candidate entity information in the text information to be processed and position weight information of each position in the text information to be processed;
[0015] Based on the first position information and the position weight information, second entity information of each candidate entity information is determined; the second entity information represents the probability of each candidate entity information appearing in the text information to be processed;
[0016] Acquire entity association information between multiple candidate entity information, and determine third entity information of each candidate entity information based on the entity association information; the third entity information represents the degree of association between each candidate entity information and other candidate entity information;
[0017] Based on the second entity information and the third entity information, first entity information of each candidate entity information is determined.
[0018] In some embodiments of the present application, the entity association information includes the number of entity associations between each candidate entity information and other candidate entity information, and determining the third entity information of each candidate entity information based on the entity association information includes:
[0019] Determine association weight information of multiple candidate entity information based on the number of entity associations;
[0020] Based on the association weight information, the initial entity information corresponding to the plurality of candidate entity information is updated to obtain updated entity information;
[0021] When the number of updates of the initial entity information does not reach the preset number, the updated entity information is determined as the initial entity information, and the step of updating the initial entity information based on the associated weight information is continued until the number of updates of the initial entity information reaches the preset number;
[0022] Based on the updated entity information, third entity information of each candidate entity information is determined.
[0023] In some embodiments of the present application, determining the first entity prompt information based on the first entity information and the plurality of candidate entity information includes:
[0024] Based on the first entity information, determine multiple target entity information from multiple candidate entity information; the multiple target entity information is entity information with higher importance among the multiple candidate entity information;
[0025] Acquire second position information of multiple target entity information in the text information to be processed;
[0026] The plurality of target entity information are sorted based on the second position information to obtain the first entity prompt information.
[0027] In some embodiments of the present application, generating target text information based on the text information to be processed and the first entity prompt information includes:
[0028] Encoding the text information to be processed and the first entity prompt information to obtain text encoding information;
[0029] The text encoding information is decoded to obtain the target text information.
[0030] In some embodiments of the present application, the text generation method is applied to a text generation model, and before generating target text information based on the text information to be processed and the first entity prompt information, the method includes:
[0031] Acquire a training text set, where the training text set includes training text and first summary information of the training text;
[0032] Extracting information from the training text to obtain second summary information of the training text;
[0033] Perform entity extraction on the second summary information to obtain second entity prompt information;
[0034] Performing masking on the training text based on the second summary information to obtain masked text of the training text;
[0035] Training a preset network model based on the mask text, the second entity prompt information, and the second summary information to obtain a pre-trained model;
[0036] Perform entity extraction on the first summary information to obtain third entity prompt information;
[0037] The pre-trained model is fine-tuned based on the training text, the third entity prompt information and the first summary information to obtain a text generation model.
[0038] In a second aspect, the present application provides a text generation device, comprising:
[0039] An information acquisition unit, used for acquiring text information to be processed;
[0040] A text processing unit, used for performing entity processing on the text information to be processed to obtain first entity prompt information;
[0041] The text generation unit is used to generate target text information based on the text information to be processed and the first entity prompt information.
[0042] In a third aspect, the present application further provides a computer device, the computer device comprising:
[0043] one or more processors;
[0044] Memory; and
[0045] One or more applications, wherein the one or more applications are stored in the memory and configured to be executed by the processor to implement any one of the text generation methods in the first aspect.
[0046] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium having a computer program stored thereon, and the computer program is loaded by a processor to execute the steps in any one of the text generation methods in the first aspect.
[0047] The beneficial effects of the present invention are as follows: entity processing is performed on text information to be processed to obtain first entity prompt information, target text information is generated based on the text information to be processed and the first entity prompt information, and the text generation task is prompted by the first entity prompt information, which can reduce the situation where the generated target text information is inconsistent with the content of the text information to be processed or the generated target text information is repeated, thereby improving the quality of the generated summary text information. BRIEF DESCRIPTION OF THE DRAWINGS
[0048] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0049] Figure 1 It is a flowchart of a text generation method provided by an embodiment of the present invention;
[0050] Figure 2 is a flowchart of a specific embodiment of the text generation method provided by an embodiment of the present invention;
[0051] Figure 3 is a flowchart of a specific embodiment of the training process of the text generation model provided by an embodiment of the present invention;
[0052] Figure 4 is a principle block diagram of a text generation device provided by an embodiment of the present invention;
[0053] Figure 5 It is a schematic diagram of the structure of a computer device provided in an embodiment of the present invention. DETAILED DESCRIPTION
[0054] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of this application.
[0055] In the description of the present application, it should be understood that the terms "center", "longitudinal", "lateral", "length", "width", "thickness", "up", "down", "front", "back", "left", "right", "vertical", "horizontal", "top", "bottom", "inside", "outside" and the like indicate positions or positional relationships based on the positions or positional relationships shown in the drawings, which are only for the convenience of describing the present application and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be understood as a limitation on the present application. In addition, the terms "first", "second", "third", "fourth", etc. are only used for descriptive purposes, and cannot be understood as indicating or implying relative importance or implicitly indicating the number of technical features indicated. Therefore, the features defined as "first", "second", "third", "fourth", etc. may explicitly or implicitly include one or more features. In the description of the present application, "multiple" means two or more, and "several" means one or more, unless otherwise clearly and specifically defined.
[0056] In this application, the word "exemplary" is used to mean "used as an example, illustration, or description." Any embodiment described in this application as "exemplary" is not necessarily to be construed as being preferred or advantageous over other embodiments. The following description is given to enable any technician in the field to implement and use the present application. In the following description, details are listed for the purpose of explanation. It should be understood that a person of ordinary skill in the art can recognize that the present application can be implemented without using these specific details. In other instances, well-known structures and processes will not be elaborated in detail to avoid obscuring the description of the present application with unnecessary details. Therefore, the present application is not intended to be limited to the embodiments shown, but is consistent with the widest scope consistent with the principles and features disclosed in the present application.
[0057] It should be noted that since the method of the embodiment of the present application is executed in a computer device, the processing objects of each computer device exist in the form of data or information. For example, time is actually time information. It can be understood that if size, quantity, position, etc. are mentioned in subsequent embodiments, they are all corresponding data for processing by the computer device. The details will not be repeated here.
[0058] The inventors have discovered through research that in existing generative summary generation tasks, long text is usually directly input into the model, and the model outputs a summary of the long text. The quality of the summary generated in this way is poor. For example, the generated summary contains content that is irrelevant or inconsistent with the long text, and the generated summary contains repeated content.
[0059] Based on this, in an embodiment of the present application, text information to be processed is obtained; entity processing is performed on the text information to be processed to obtain first entity prompt information; based on the text information to be processed and the first entity prompt information, target text information is generated, and the text generation task is prompted by the first entity prompt information, which can reduce the situation where the generated target text information is inconsistent with the content of the text information to be processed or the generated target text information is repeated, thereby improving the quality of the generated summary text information.
[0060] The content of this application is further illustrated by describing embodiments in conjunction with the accompanying drawings.
[0061] This embodiment provides a text generation method, such as Figure 1 As shown in , the method includes:
[0062] Step S10: Obtain text information to be processed.
[0063] The text information to be processed is the text that needs to be processed in the text generation task. The text information to be processed can be text data input by the user through a computer device (for example, a smart phone), or text data obtained by converting the collected user conversations through a computer device. For example, the text information to be processed can be a news report on a hot event input by the user through a mobile phone.
[0064] Step S20: perform entity processing on the text information to be processed to obtain first entity prompt information.
[0065] Entity processing includes entity selection and entity sorting of the text information to be processed. The first entity prompt information is an entity chain obtained by entity processing the text information to be processed. The first entity prompt information includes multiple entities with high importance in the text information to be processed. As prompt information in the text generation task, it can reduce the situation where the generated text is irrelevant to the text information to be processed or duplicate content appears in the generated text, thereby improving the quality of the generated text.
[0066] In a specific implementation, Figure 2 As shown, step S20 includes:
[0067] Step S21, extracting entities from the text information to be processed to obtain multiple candidate entity information of the text information to be processed;
[0068] Step S22: determining first entity information of each candidate entity information; the first entity information represents the importance of each candidate entity information;
[0069] Step S23: Determine first entity prompt information based on the first entity information and multiple candidate entity information.
[0070] Entity, also known as named entity, refers to names of people, organizations, places, and all other entities identified by names. Multiple candidate entity information is entity information extracted from the text information to be processed. For example, the BosonNLP tool can be used to extract entities from the text information to be processed. The BosonNLP tool has a high recognition ability for entity information such as time, place, name of a person, organization name, company name, product name, position, etc. Using the BosonNLP tool to extract entities from the text information to be processed can improve the accuracy of the extracted multiple candidate entity information.
[0071] The first entity information is used to characterize the importance of each candidate entity information. In order to select entity information with higher importance from multiple candidate entity information to form the first entity prompt information, when performing entity processing on the text information to be processed, this embodiment first performs entity extraction on the text information to be processed, extracts multiple candidate entity information from the text information to be processed, and then determines the first entity information of each candidate entity information. Finally, based on the first entity information and the multiple candidate entity information, the first entity prompt information is determined.
[0072] In a specific implementation, continue to refer to Figure 2 As shown, step S22 includes:
[0073] Step S221, obtaining first position information of each candidate entity information in the text information to be processed and position weight information of each position in the text information to be processed;
[0074] Step S222: determining second entity information of each candidate entity information based on the first position information and the position weight information; the second entity information represents the probability of each candidate entity information appearing in the text information to be processed;
[0075] Step S223: Acquire entity association information between multiple candidate entity information, and determine third entity information of each candidate entity information based on the entity association information; the third entity information represents the degree of association between each candidate entity information and other candidate entity information;
[0076] Step S224: Determine the first entity information of each candidate entity information based on the second entity information and the third entity information.
[0077] The first position information is the position information of each candidate entity information in the text information to be processed. For example, in "In the winter of 1983, the deaf-mute homeless man Jiang Facai was taken home by villager Jiang Yinhong. 32 years later, Jiang Yinhong died of a sudden cerebral hemorrhage", the entity "Jiang Facai" is at the 11th to 13th character position, and the entity "Jiang Yinhong" is at the 17th to 19th character position.
[0078] Taking into account the different importance of entities at different positions in the text information to be processed, for example, the importance of the first sentence, first paragraph, last sentence, last paragraph and other entities is first sentence > first paragraph = last sentence > last paragraph > others, this embodiment pre-sets different position weight information for different positions in the text information to be processed. For example, the position weight information of the first sentence, first paragraph, last sentence, last paragraph and other positions is 4, 3, 3, 2, 1 respectively.
[0079] The second entity information is the entity information of each candidate entity information determined based on the first position information and the position weight information. The second entity information can characterize the probability of each candidate entity information appearing in the text information to be processed. In a specific implementation, the step of determining the second entity information of each candidate entity information based on the first position information and the position weight information can specifically include: determining the entity weight information of each candidate entity information based on the first position information and the position weight information, the entity weight information being the sum of the position weight information of the position where each candidate entity information appears in the text information to be processed; determining the second entity information of each candidate entity information based on the entity weight information. Among them, the formula for determining the second entity information is: P i Represents the second entity information of the i-th candidate entity information, l i represents the entity weight information of the i-th candidate entity information, m represents the number of multiple candidate entity information, l j Represents the entity weight information of the j-th candidate entity information.
[0080] The entity association information is the association information between multiple candidate entity information, and the third entity information is the entity information determined based on the entity association information. Considering that entities that have more associations with other entities are more important and entities that are associated with more important entities are more important, this embodiment obtains the entity association information between multiple candidate entity information, determines the third entity information of each candidate entity information based on the entity association information, and then determines the first entity information of each candidate entity information based on the second entity information and the third entity information. The calculation formula of the first entity information is: S(v i ) represents the i-th candidate entity information v i The first entity information, P i Represents the second entity information of the i-th candidate entity information, s(v i ) represents the i-th candidate entity information v i The third entity information, m represents the number of candidate entity information, P j Represents the second entity information of the jth candidate entity information, s(v j ) represents the jth candidate entity information v j The third entity information.
[0081] In a specific implementation, the entity association information includes the number of entity associations between each candidate entity information and other candidate entity information. The step of determining the third entity information of each candidate entity information based on the entity association information in step S223 includes:
[0082] Step S2231, determining association weight information of multiple candidate entity information based on the number of entity associations;
[0083] Step S2232: updating the initial entity information corresponding to the plurality of candidate entity information based on the association weight information to obtain updated entity information;
[0084] Step S2233: when the number of updates of the initial entity information does not reach the preset number, the updated entity information is determined as the initial entity information, and the step of updating the initial entity information based on the associated weight information is continued until the number of updates of the initial entity information reaches the preset number;
[0085] Step S2234: Determine the third entity information of each candidate entity information based on the updated entity information.
[0086] In a specific implementation, the entity association information includes the number of entity associations between each candidate entity information and other candidate entity information. The number of entity associations refers to the number of times each candidate entity information appears in the same sentence with other candidate entity information. For example, candidate entity information A and candidate entity information B appear simultaneously in the second sentence, the fifth sentence, and the tenth sentence in the text information to be processed, respectively. Then the number of entity associations between candidate entity information A and candidate entity information B is 3.
[0087] The association weight information is an association weight matrix of multiple candidate entity information. The association weight information is an m×m matrix, where m represents the number of multiple candidate entity information. The association weight information includes m×m elements, and the element w ij Represents the i-th candidate entity information v i With the jth candidate entity information v j The association weight of the candidate entity information u i and candidate entity information v j The association weight is the candidate entity information v i and candidate entity information v j The ratio of the number of associations between the entity association times and the maximum number of associations between the multiple candidate entity information. The step of determining the association weight information of the multiple candidate entity information based on the number of entity associations may specifically include: determining the maximum number of associations between the multiple candidate entity information based on the number of entity associations; determining the association weights between the multiple candidate entity information based on the number of entity associations and the maximum number of associations; and determining the association weight information of the multiple candidate entity information based on the association weights.
[0088] The initial entity information is an initial matrix constructed based on the number of candidate entity information. The initial entity information is an m×1 matrix. The initial entity information includes m elements, and the value of each element is The initial entity information can be expressed as: m represents the number of candidate entity information.
[0089] The updated entity information is the entity information obtained by updating the initial entity information based on the associated weight information. In order to more accurately measure the importance of multiple candidate entity information, this embodiment constructs the initial entity information corresponding to the multiple candidate entity information, and updates the initial entity information based on the associated weight information to obtain the updated entity information. When the number of updates of the initial entity information does not reach the preset number, the updated entity information is determined as the initial entity information, and the step of updating the initial entity information based on the associated weight information is continued until the number of updates of the initial entity information reaches the preset number. The updating process of the initial entity information can be expressed as: S n represents the initial entity information after the nth update, S n-1 represents the initial entity information after the n-1th update, d represents the preset damping factor, and W represents the associated weight information.
[0090] The updated entity information obtained after updating the initial entity information a preset number of times includes m elements, and the m elements are the third entity information of the m candidate entity information. Therefore, after the initial entity information is updated, the third entity information of each candidate entity information can be determined based on the updated entity information.
[0091] In a specific implementation, continue to refer to Figure 2 As shown, step S23 includes:
[0092] Step S231: Based on the first entity information, determine multiple target entity information from multiple candidate entity information; the multiple target entity information is entity information with higher importance among the multiple candidate entity information;
[0093] Step S232, obtaining second position information of multiple target entity information in the text information to be processed;
[0094] Step S233: sort the multiple target entity information based on the second location information to obtain first entity prompt information.
[0095] The multiple target entity information is the entity information with the highest importance among the multiple candidate entity information. For example, the multiple target entity information is the k entity information with the highest importance among the multiple candidate entity information. The second position information is the position where the multiple target entity information appears in the text information to be processed. Based on the second position information, the order in which the multiple target entity information appears in the text information to be processed can be determined.
[0096] In order to increase the diversity of the generated target text information, after determining multiple target entity information with a higher importance from multiple candidate entity information, this embodiment further obtains the second position information of the multiple target entity information in the text information to be processed, and then determines the order in which the multiple target entity information appears in the text to be processed based on the second position information, and sorts the multiple target entity information based on the order in which the multiple target entity information appears in the text to be processed, thereby obtaining the first entity prompt information. For example, target text A appears for the first time in the first sentence of the text information to be processed, target text B appears for the first time in the fifth sentence of the text information to be processed, and target text C appears for the first time in the third sentence of the text information to be processed, then the sorting order of target text A, target text B, and target text C is target text A, target text C, and target text B.
[0097] S30: Generate target text information based on the text information to be processed and the first entity prompt information.
[0098] The target text information is a summary text of the text information to be processed generated based on the first entity prompt information. The first entity prompt information serves as prompt information for the text generation task. Generating the target text information based on the first entity prompt information can reduce the situation where the generated text is irrelevant to the text information to be processed or repeated content appears in the generated text, thereby improving the quality of the generated summary text.
[0099] In a specific implementation, the text generation method is applied to a text generation model, and the step of generating target text information based on the text information to be processed and the first entity prompt information may specifically include: inputting the text information to be processed and the first entity prompt information into the text generation model, and outputting the target text information through the text generation model.
[0100] Furthermore, the text generation model includes an encoder and a decoder, wherein the encoder may be an encoder of a Transformer structure, and the decoder may be a decoder of a Transformer structure. The steps of inputting the text information to be processed and the first entity prompt information into the text generation model, and outputting the target text information through the text generation model may specifically include: inputting the text information to be processed and the first entity prompt information into the encoder, encoding the text information to be processed and the first entity prompt information through the encoder to obtain text encoding information; inputting the text encoding information into the decoder, and decoding the text encoding information through the decoder to obtain the target text information.
[0101] In a specific implementation, Figure 3 As shown, before step S30, it includes:
[0102] Step S41, obtaining a training text set, wherein the training text set includes training text and first summary information of the training text;
[0103] Step S42: extracting information from the training text to obtain second summary information of the training text;
[0104] Step S43: extracting entities from the second summary information to obtain second entity prompt information;
[0105] Step S44, masking the training text based on the second summary information to obtain a masked text of the training text;
[0106] Step S45: training a preset network model based on the mask text, the second entity prompt information and the second summary information to obtain a pre-trained model;
[0107] Step S46: extract entity from the first summary information to obtain third entity prompt information;
[0108] Step S47: fine-tune the pre-trained model based on the training text, the third entity prompt information and the first summary information to obtain a text generation model.
[0109] The training text set is a set of texts preset for training the preset network model, and the training text set includes the training text and the first summary information of the training text. The preset network model has the same structure as the text generation model, and the difference between the preset network model and the text generation model is that the model parameters of the preset network model are the initial model parameters, and the model parameters of the text generation model are the trained model parameters.
[0110] The second summary information is the summary information extracted from the training text, which is formed by the fusion of multiple sentences with the highest importance in the training text. The second entity prompt information is an entity chain composed of entities extracted from the second summary information. For example, the second summary information "In the winter of 1983, the deaf and mute homeless man Jiang Facai was taken home by villager Jiang Yinhong. 32 years have passed, and Jiang Yinhong has died of a sudden cerebral hemorrhage. But Jiang Yinhong's family's affection for Jiang Facai has not changed. To outsiders, they are completely a family, and it is impossible to tell that Jiang Facai was "picked up." " can be extracted to obtain the second entity prompt information "1983 | Winter | Jiang Facai | Jiang Yinhong | 32 years | Jiang Yinhong | Cerebral hemorrhage ||| Jiang Yinhong | Jiang Facai | Jiang Facai", and the entities of the same sentence in the second entity prompt information are separated by the "|" symbol, and the entities of different sentences are separated by the '|||' symbol.
[0111] The masked text is the text obtained by masking the training text based on the second summary information. In order to improve the generalization performance of the text generation model, when masking the training text, in addition to masking the second summary information in the training text, the words of a preset proportion in the training text can also be masked. For example, the second summary information in the training text is replaced by [MASK1], and 10% of the words are randomly selected to be replaced by [MASK2]. For example, the masked text is "[MASK1] In the boiling [MASK2] online public opinion [MASK2], the author [MASK2] published [MASK2] two concepts of the Chinese people...", [MASK1] represents the second summary information, and [MASK2] represents the randomly masked words.
[0112] The pre-trained model is a trained model obtained by training a preset network model based on masked text, second entity prompt information and second summary information. The process of training a preset network model based on masked text, second entity prompt information and second summary information may specifically include: inputting masked text and second entity prompt information into the preset network model, and outputting the first predicted summary information of the masked text through the preset network model; determining the first loss value based on the first predicted summary information, the second summary information and the loss function of the preset network model; when the first loss value does not meet the preset first condition, correcting the model parameters of the preset network model based on the preset first parameter learning rate, and continuing to execute the step of inputting the masked text and the second entity prompt information into the preset network model until the first loss value meets the first condition, thereby obtaining the pre-trained model. Among them, the loss function of the preset network model can adopt the maximum likelihood estimation (MLE) function.
[0113] The third entity prompt information is an entity chain composed of entities extracted from the first summary information. For example, the BosonNLP tool can be used to extract entities from the first summary information to obtain the first entity chain information. In order to enable the pre-trained model to better match downstream tasks, after obtaining the pre-trained model, this embodiment further fine-tunes the pre-trained model based on the training text, the third entity prompt information and the first summary information to obtain a text generation model.
[0114] In a specific implementation, the step of fine-tuning the pre-trained model based on the training text, the third entity prompt information and the first summary information may specifically include: inputting the training text and the third entity prompt information into the pre-trained model, and outputting the second predicted summary information of the training text through the pre-trained model; determining the second loss value based on the second predicted summary information, the first summary information and the loss function of the pre-trained model; when the second loss value does not meet the preset second condition, correcting the model parameters of the pre-trained model based on the preset second parameter learning rate, and continuing to perform the step of inputting the training text and the third entity prompt information into the pre-trained model until the second loss value meets the second condition, thereby obtaining the text generation model. Among them, the loss function of the pre-trained model can adopt the maximum likelihood estimation (MLE) function.
[0115] In a specific implementation, step S42 includes:
[0116] Step S421, obtaining the third position information of each sentence in the training text;
[0117] Step S422: determining the position sequence information of each sentence based on the third position information;
[0118] Step S423: determining the first text information of each sentence based on the position sequence information;
[0119] Step S424, determining the vocabulary overlap information of each sentence in the training text, where the vocabulary overlap information represents the degree of vocabulary overlap between each sentence and other sentences in the training text;
[0120] Step S425: determining second text information of each sentence based on the first text information and the vocabulary overlap information; the second text information represents the importance of each sentence;
[0121] Step S426: determining a plurality of target sentences based on the second text information; the plurality of target sentences are sentences with a high degree of importance in the training text;
[0122] Step S427: fuse multiple target sentences to obtain second summary information of the training text.
[0123] The training text includes multiple sentences. The third position information is the position information of each sentence in the training text. The position order information represents the position importance order of each sentence. The i-th sentence x in the training text i The position order information of can be expressed as j, where j∈(1,n) and n represents the number of sentences in the training text.
[0124] Considering that sentences at different positions have different importances, for example, sentences at the beginning and end of a text are relatively more important, and sentences at the beginning are more important than sentences at the end, this embodiment obtains the third position information of each sentence in the training text, reorders multiple sentences based on the third position information, obtains the sorting results of multiple sentences, and then determines the position order information of each sentence based on the sorting results of multiple sentences. For example, the first 15% of the sentences in the training text are taken as the beginning part, and the last 10% of the sentences in the training text are taken as the ending part. The beginning part is first sorted according to the order in the training text, and then the ending part is sorted in reverse order. The sorting result is sentence A, sentence B, sentence O, sentence C, sentence D, ..., sentence N, then the position order information of sentence A, sentence B, sentence O, sentence C, sentence D, ..., sentence N is 1, 2, ..., 14, 15 respectively.
[0125] The first text information represents the position importance of each sentence in the training text. The first text information is determined based on the position order information of each sentence and the number of sentences in the training text. The calculation formula of the first text information is: s i ′ represents the first text information of the i-th sentence in the training text, n represents the number of sentences in the training text, j represents the position order information of the i-th sentence in the training text, and λ represents the preset strength parameter.
[0126] The vocabulary overlap information represents the degree of vocabulary overlap between each sentence in the training text and other sentences in the training text. The vocabulary overlap information of the i-th sentence in the training text can be expressed as: s i ″ represents the vocabulary overlap information of the i-th sentence in the training text, x i represents the i-th sentence in the training text, and n represents the number of sentences in the training text.
[0127] The second text information represents the importance of each sentence in the training text. The second text information is determined based on the first text information and the vocabulary overlap information. The calculation formula of the second text information is: i =γs i ′+(1-γ)s i ″, si represents the second text information of the i-th sentence in the training text, s i′ represents the first text information of the i-th sentence in the training text, s i ″ represents the vocabulary overlap information of the ith sentence in the training text, and γ represents the pre-set position importance weight of the ith sentence.
[0128] The multiple target sentences are sentences with the highest importance in the training text. After determining the second text information of each sentence in the training text, multiple sentences with the highest importance can be selected from the training text as target sentences based on the second text information, and the multiple target sentences can be spliced to obtain the second summary information of the training text. For example, based on the second text information, the importance of the sentences in the training text is determined as sentence D>sentence A>sentence F>sentence B>sentence C>sentence E, and the three most important sentences, namely sentence D, sentence A and sentence B, are selected from the training text for splicing to obtain the second summary information of the training text.
[0129] In order to better implement the text generation method in the embodiment of the present application, based on the text generation method, the embodiment of the present application also provides a text generation device, such as Figure 4 As shown, the text generation device includes:
[0130] The information acquisition unit 610 is used to acquire the text information to be processed;
[0131] A text processing unit 620, configured to perform entity processing on the text information to be processed to obtain first entity prompt information;
[0132] The text generation unit 630 is used to generate target text information based on the text information to be processed and the first entity prompt information.
[0133] In an embodiment of the present application, entity processing is performed on the text information to be processed to obtain first entity prompt information, target text information is generated based on the text information to be processed and the first entity prompt information, and the text generation task is prompted by the first entity prompt information. This can reduce the situation where the generated target text information is inconsistent with the content of the text information to be processed or the generated target text information is repeated, thereby improving the quality of the generated summary text information.
[0134] In some embodiments of the present application, the text processing unit 620 is specifically used for:
[0135] Perform entity extraction on the text information to be processed to obtain multiple candidate entity information of the text information to be processed;
[0136] Determine first entity information of each candidate entity information; the first entity information represents the importance of each candidate entity information;
[0137] First entity prompt information is determined based on the first entity information and the plurality of candidate entity information.
[0138] In some embodiments of the present application, the text processing unit 620 is further configured to:
[0139] Obtaining first position information of each candidate entity information in the text information to be processed and position weight information of each position in the text information to be processed;
[0140] Based on the first position information and the position weight information, second entity information of each candidate entity information is determined; the second entity information represents the probability of each candidate entity information appearing in the text information to be processed;
[0141] Acquire entity association information between multiple candidate entity information, and determine third entity information of each candidate entity information based on the entity association information; the third entity information represents the degree of association between each candidate entity information and other candidate entity information;
[0142] Based on the second entity information and the third entity information, first entity information of each candidate entity information is determined.
[0143] In some embodiments of the present application, the entity association information includes the number of entity associations of each candidate entity information with other candidate entity information. The text processing unit 620 is further configured to:
[0144] Determine association weight information of multiple candidate entity information based on the number of entity associations;
[0145] Based on the association weight information, the initial entity information corresponding to the plurality of candidate entity information is updated to obtain updated entity information;
[0146] When the number of updates of the initial entity information does not reach the preset number, the updated entity information is determined as the initial entity information, and the step of updating the initial entity information based on the associated weight information is continued until the number of updates of the initial entity information reaches the preset number;
[0147] Based on the updated entity information, third entity information of each candidate entity information is determined.
[0148] In some embodiments of the present application, the text processing unit 620 is further configured to:
[0149] Based on the first entity information, determine multiple target entity information from multiple candidate entity information; the multiple target entity information is entity information with higher importance among the multiple candidate entity information;
[0150] Acquire second position information of multiple target entity information in the text information to be processed;
[0151] The plurality of target entity information are sorted based on the second position information to obtain the first entity prompt information.
[0152] In some embodiments of the present application, the text generation unit 630 is specifically used to:
[0153] Encoding the text information to be processed and the first entity prompt information to obtain text encoding information;
[0154] The text encoding information is decoded to obtain the target text information.
[0155] In some embodiments of the present application, the text generation method is applied to a text generation model, and the text processing device further includes:
[0156] A text acquisition unit, used to acquire a training text set, wherein the training text set includes a training text and first summary information of the training text;
[0157] A first extraction unit is used to extract information from the training text to obtain second summary information of the training text;
[0158] A second extraction unit, configured to extract entities from the second summary information to obtain second entity prompt information;
[0159] A mask processing unit, configured to perform mask processing on the training text based on the second summary information to obtain a masked text of the training text;
[0160] A model training unit, used to train a preset network model based on the mask text, the second entity prompt information and the second summary information to obtain a pre-trained model;
[0161] A third extraction unit is used to extract entities from the first summary information to obtain third entity prompt information;
[0162] The model fine-tuning unit is used to fine-tune the pre-trained model based on the training text, the third entity prompt information and the first summary information to obtain a text generation model.
[0163] The embodiment of the present application further provides a computer device, which integrates any text generation device provided in the embodiment of the present application, and the computer device includes:
[0164] one or more processors;
[0165] Memory; and
[0166] One or more applications, wherein the one or more applications are stored in a memory and are configured to execute, by a processor, the steps of the text generation method in any of the above-mentioned text generation method embodiments.
[0167] The present application also provides a computer device that integrates any of the text generation devices provided in the present application. Figure 5As shown, it shows a schematic diagram of the structure of the computer device involved in the embodiment of the present application, specifically:
[0168] The computer device may include one or more processing core processors 801, one or more computer-readable storage media memories 802, a power supply 803, an input unit 804 and other components. Those skilled in the art will appreciate that Figure 5 The computer device structure shown in the figure does not constitute a limitation on the computer device, and may include more or fewer components than shown in the figure, or combine certain components, or arrange the components differently. Among them:
[0169] The processor 801 is the control center of the computer device. It uses various interfaces and lines to connect various parts of the entire computer device. By running or executing software programs and / or modules stored in the memory 802 and calling data stored in the memory 802, it executes various functions of the computer device and processes data, thereby monitoring the computer device as a whole. Optionally, the processor 801 may include one or more processing cores; preferably, the processor 801 may integrate an application processor and a modem processor, wherein the application processor mainly processes the operating system, user interface, and application programs, and the modem processor mainly processes wireless communications. It is understandable that the above-mentioned modem processor may not be integrated into the processor 801.
[0170] The memory 802 can be used to store software programs and modules. The processor 801 executes various functional applications and data processing by running the software programs and modules stored in the memory 802. The memory 802 may mainly include a program storage area and a data storage area, wherein the program storage area may store an operating system, an application required for at least one function (such as a sound playback function, an image playback function, etc.), etc.; the data storage area may store data created according to the use of the computer device, etc. In addition, the memory 802 may include a high-speed random access memory, and may also include a non-volatile memory, such as at least one disk storage device, a flash memory device, or other volatile solid-state storage devices. Accordingly, the memory 802 may also include a memory controller to provide the processor 801 with access to the memory 802.
[0171] The computer device also includes a power supply 803 for supplying power to each component. Preferably, the power supply 803 can be logically connected to the processor 801 through a power management system, so that the power management system can manage charging, discharging, power consumption and other functions. The power supply 803 can also include one or more DC or AC power supplies, recharging systems, power failure detection circuits, power converters or inverters, power status indicators and other arbitrary components.
[0172] The computer device may further include an input unit 804, which may be used to receive input digital or character information and generate keyboard, mouse, joystick, optical or trackball signal input related to user settings and function control.
[0173] Although not shown, the computer device may further include a display unit, etc., which will not be described in detail herein. Specifically in this embodiment, the processor 801 in the computer device will load the executable files corresponding to the processes of one or more application programs into the memory 802 according to the following instructions, and the processor 801 will run the application programs stored in the memory 802, thereby realizing various functions, as follows:
[0174] Get the text information to be processed;
[0175] Perform entity processing on the text information to be processed to obtain first entity prompt information;
[0176] Target text information is generated based on the text information to be processed and the first entity prompt information.
[0177] A person of ordinary skill in the art will appreciate that all or part of the steps in the various methods of the above embodiments may be completed by instructions, or by controlling related hardware through instructions. The instructions may be stored in a computer-readable storage medium and loaded and executed by a processor.
[0178] To this end, the embodiment of the present application provides a computer-readable storage medium, which may include: a read-only memory (ROM), a random access memory (RAM), a disk or an optical disk, etc. A computer program is stored thereon, and the computer program is loaded by a processor to execute the steps in any text generation method provided in the embodiment of the present application. For example, the computer program is loaded by a processor to execute the following steps:
[0179] Get the text information to be processed;
[0180] Perform entity processing on the text information to be processed to obtain first entity prompt information;
[0181] Target text information is generated based on the text information to be processed and the first entity prompt information.
[0182] In the above embodiments, the description of each embodiment has its own emphasis. For parts that are not described in detail in a certain embodiment, please refer to the detailed description of other embodiments above, and will not be repeated here.
[0183] In specific implementation, the above units or structures can be implemented as independent entities, or can be arbitrarily combined to be implemented as the same or several entities. The specific implementation of the above units or structures can refer to the previous method embodiments, which will not be repeated here.
[0184] The specific implementation of the above operations can be found in the previous embodiments, which will not be described in detail here.
[0185] The above is a detailed introduction to a text generation method, device, computer equipment and computer-readable storage medium provided in the embodiments of the present application. Specific examples are used in this article to illustrate the principles and implementation methods of the present application. The description of the above embodiments is only used to help understand the method of the present application and its core idea; at the same time, for technical personnel in this field, according to the ideas of the present application, there will be changes in the specific implementation methods and application scopes. In summary, the content of this specification should not be understood as a limitation on the present application.
Claims
1. A text generation method, characterized in that: include: Get the text information to be processed; Performing entity processing on the text information to be processed to obtain first entity prompt information; Generate target text information based on the text information to be processed and the first entity prompt information.
2. The method according to claim 1, characterized in that The performing entity processing on the text information to be processed to obtain first entity prompt information includes: Performing entity extraction on the text information to be processed to obtain multiple candidate entity information of the text information to be processed; Determine first entity information of each candidate entity information; the first entity information represents the importance of each candidate entity information; First entity prompt information is determined based on the first entity information and the plurality of candidate entity information.
3. The method according to claim 2, characterized in that The determining of the first entity information of each candidate entity information includes: Obtaining first position information of each candidate entity information in the text information to be processed and position weight information of each position in the text information to be processed; Based on the first position information and the position weight information, determining second entity information of each candidate entity information; the second entity information represents the probability of each candidate entity information appearing in the text information to be processed; Acquire entity association information between a plurality of candidate entity information, and determine third entity information of each candidate entity information based on the entity association information; the third entity information represents the degree of association between each candidate entity information and other candidate entity information; Based on the second entity information and the third entity information, first entity information of each candidate entity information is determined.
4. The method according to claim 3, characterized in that The entity association information includes the number of entity associations between each candidate entity information and other candidate entity information, and determining the third entity information of each candidate entity information based on the entity association information includes: Determining association weight information of a plurality of candidate entity information based on the entity association times; Based on the association weight information, the initial entity information corresponding to the plurality of candidate entity information is updated to obtain updated entity information; When the number of updates of the initial entity information does not reach the preset number, determining the updated entity information as the initial entity information, and continuing to perform the step of updating the initial entity information based on the association weight information until the number of updates of the initial entity information reaches the preset number; Based on the updated entity information, third entity information of each candidate entity information is determined.
5. The method according to claim 2, characterized in that: The determining the first entity prompt information based on the first entity information and the plurality of candidate entity information includes: Based on the first entity information, determine multiple target entity information from the multiple candidate entity information; the multiple target entity information is the entity information with the highest importance among the multiple candidate entity information; Acquire second position information of a plurality of target entity information in the text information to be processed; The plurality of target entity information are sorted based on the second position information to obtain first entity prompt information.
6. The method according to claim 1, characterized in that The generating target text information based on the text information to be processed and the first entity prompt information includes: Encoding the text information to be processed and the first entity prompt information to obtain text encoding information; The text encoding information is decoded to obtain target text information.
7. The method according to claim 1, characterized in that The text generation method is applied to a text generation model, and before generating target text information based on the text information to be processed and the first entity prompt information, it includes: Acquire a training text set, wherein the training text set includes training text and first summary information of the training text; Extracting information from the training text to obtain second summary information of the training text; Performing entity extraction on the second summary information to obtain second entity prompt information; Performing masking on the training text based on the second summary information to obtain masked text of the training text; Training a preset network model based on the mask text, the second entity prompt information and the second summary information to obtain a pre-trained model; Performing entity extraction on the first summary information to obtain third entity prompt information; The pre-training model is fine-tuned based on the training text, the third entity prompt information and the first summary information to obtain a text generation model.
8. A text generation device, characterized in that: include: An information acquisition unit, used for acquiring text information to be processed; A text processing unit, used for performing entity processing on the text information to be processed to obtain first entity prompt information; A text generation unit is used to generate target text information based on the text information to be processed and the first entity prompt information.
9. A computer device, characterized in that: The computer device comprises: one or more processors; Memory; and One or more application programs, wherein the one or more application programs are stored in the memory and configured to be executed by the processor to implement the text generation method according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that: A computer program is stored thereon, and the computer program is loaded by a processor to execute the steps in the text generation method according to any one of claims 1 to 7.