Plot text processing method and device

By introducing plot categories into the training samples and extracting target training samples to train the summary generation model, the problem of inaccurate generation by traditional models is solved, and accurate generation of plot summaries is achieved.

CN118690011BActive Publication Date: 2025-11-21TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410962931.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-07-17
Publication Date
2025-11-21
Estimated Expiration
2044-07-17

AI Technical Summary

Technical Problem

Traditional summary generation models suffer from inaccurate plot summaries and poor generalization when the training data is biased, making them prone to producing erroneous results.

Method used

By acquiring multiple training samples, determining the plot category according to the classification method, and extracting target training samples from them, a summary generation model is trained. By introducing joint learning of plot categories, accurate plot summary generation can be achieved.

Benefits of technology

The generation of plot summaries has been improved across different plot categories, reducing the possibility of misunderstanding and generating accurate plot summaries.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118690011B_ABST
    Figure CN118690011B_ABST
Patent Text Reader

Abstract

The application relates to a plot text processing method and device, computer equipment, a storage medium and a computer program product. The method comprises the following steps: obtaining a plurality of training samples; obtaining the number of classification modes used for determining the plot categories to which the plurality of training samples belong, and obtaining the number of plot categories configured for the training samples used for a batch of training; determining the number of plot categories extracted from each plot category under each classification mode of the plurality of training samples according to the number of classification modes and the number of plot categories; for each classification mode, a target plot category is extracted from the plot categories under the classification mode according to the number of plot categories extracted, and a target training sample is extracted from the training samples containing the target plot category, so as to obtain the training samples used for a batch of training; and training an abstract generation model according to the training samples used for a batch of training. The method can generate an accurate plot abstract.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of computer, in particular to a plot text processing method and device, computer equipment, storage medium and computer program product. BACKGROUND

[0002] With the development of artificial intelligence technology, natural language processing technology is constantly developing, and the summary generation model based on natural language processing technology has been widely applied. The summary generation model aims to obtain the key information of the text and generate a short summary containing the key information. The summary generation model can be applied to the generation of plot summaries corresponding to plot texts. It can be understood that generating plot summaries can help quickly understand the plot content.

[0003] In the traditional method, on the basis of training a summary generation model by using training samples containing plot summaries and plot texts, by inputting the plot text of a certain game into the summary generation model, the plot summary of the input plot text generated by the summary generation model can be obtained.

[0004] However, although the summary generation model in the traditional method can quickly generate plot summaries, the training data when training the summary generation model is biased, which leads to poor generalization of the summary generation model and inaccurate results when predicting, and there is a situation that the generated plot summary is inaccurate. SUMMARY

[0005] Therefore, it is necessary to provide a plot text processing method, device, computer equipment, computer readable storage medium and computer program product capable of generating accurate plot summaries in view of the above technical problems.

[0006] In a first aspect, the present application provides a plot text processing method. The method comprises:

[0007] Obtaining a plurality of training samples, the training samples comprising plot texts, plot summaries and at least one plot category to which the plot texts belong, each plot category being obtained by classifying the plot texts according to a classification manner;

[0008] Obtaining the number of classification manners used to determine the plot categories to which the plurality of training samples belong, and obtaining the number of plot categories configured for a batch of training samples used for training;

[0009] According to the number of classification manners and the number of plot categories, determining the number of plot categories to be extracted from each classification manner of the plurality of training samples;

[0010] For each classification manner, a target plot category is extracted from plot categories of the plurality of training samples under the classification manner according to the plot category extraction quantity, and a target training sample is extracted from the training sample containing the target plot category, to obtain training samples used by a batch of training;

[0011] An abstract generation model is trained according to the training samples used by the batch of training, and the trained abstract generation model is used to output a corresponding plot abstract of the to-be-processed plot text based on input of the to-be-processed plot text.

[0012] In a second aspect, the present application further provides a plot text processing device. The device comprises:

[0013] A sample acquisition module is configured to acquire a plurality of training samples, wherein each training sample comprises plot text, a plot abstract, and at least one plot category to which the plot text belongs, and each plot category is obtained by classifying the plot text according to a classification manner;

[0014] A quantity acquisition module is configured to acquire a quantity of classification manners used to determine plot categories to which the plurality of training samples belong, and acquire a plot category quantity configured for training samples used by a batch of training;

[0015] An extraction quantity determination module is configured to determine, according to the quantity of classification manners and the plot category quantity, a plot category extraction quantity of extracting plot categories from plot categories of the plurality of training samples under each classification manner;

[0016] A training sample extraction module is configured to, for each classification manner, extract a target plot category from plot categories of the plurality of training samples under the classification manner according to the plot category extraction quantity, and extract a target training sample from the training sample containing the target plot category, to obtain training samples used by a batch of training;

[0017] A training module is configured to train an abstract generation model according to the training samples used by the batch of training, and the trained abstract generation model is used to output a corresponding plot abstract of to-be-processed plot text based on input of the to-be-processed plot text.

[0018] In a third aspect, the present application further provides a computer device. The computer device comprises a memory and a processor, wherein the memory stores a computer program, and the processor implements the following steps when executing the computer program:

[0019] obtaining a plurality of training samples, wherein each of the training samples comprises a plot text, a plot abstract, and at least one plot category to which the plot text belongs, and each of the plot categories is obtained by classifying the plot text according to a classification manner;

[0020] obtaining a number of classification manners used for determining the plot categories to which the plurality of training samples belong, and obtaining a number of plot categories configured for training samples used for a batch training;

[0021] determining, according to the number of classification manners and the number of plot categories, a number of plot categories to be extracted from the plot categories under each of the classification manners of the plurality of training samples;

[0022] for each of the classification manners, extracting, according to the number of plot categories to be extracted, a target plot category from the plot categories under the classification manner of the plurality of training samples, and extracting a target training sample from the training samples containing the target plot category, to obtain the training samples used for the batch training;

[0023] training, according to the training samples used for the batch training, an abstract generation model, wherein the trained abstract generation model is configured to output a plot abstract corresponding to a plot text to be processed based on the plot text to be processed.

[0024] In a fourth aspect, the present application further provides a computer readable storage medium. The computer readable storage medium has a computer program stored thereon, and the computer program is executed by a processor to implement the following steps:

[0025] obtaining a plurality of training samples, wherein each of the training samples comprises a plot text, a plot abstract, and at least one plot category to which the plot text belongs, and each of the plot categories is obtained by classifying the plot text according to a classification manner;

[0026] obtaining a number of classification manners used for determining the plot categories to which the plurality of training samples belong, and obtaining a number of plot categories configured for training samples used for a batch training;

[0027] determining, according to the number of classification manners and the number of plot categories, a number of plot categories to be extracted from the plot categories under each of the classification manners of the plurality of training samples;

[0028] for each of the classification manners, extracting, according to the number of plot categories to be extracted, a target plot category from the plot categories under the classification manner of the plurality of training samples, and extracting a target training sample from the training samples containing the target plot category, to obtain the training samples used for the batch training;

[0029] According to the training samples used for the batch training, a summary generation model is trained; the trained summary generation model is used to output a corresponding plot summary of the to-be-processed plot text based on the input to-be-processed plot text.

[0030] In a fifth aspect, the present application further provides a computer program product. The computer program product comprises a computer program which, when executed by a processor, implements the following steps:

[0031] Obtaining a plurality of training samples, wherein each training sample comprises plot text, a plot summary, and at least one plot category to which the plot text belongs, and each plot category is obtained by classifying the plot text according to a classification manner;

[0032] Obtaining the number of classification manners used for determining the plot categories to which the plurality of training samples belong, and obtaining the number of plot categories configured for the training samples used for a batch training;

[0033] According to the number of classification manners and the number of plot categories, determining the number of plot categories to be extracted from the plot categories under each classification manner of the plurality of training samples;

[0034] For each classification manner, according to the number of plot categories to be extracted, extracting target plot categories from the plot categories under the classification manner of the plurality of training samples, and extracting target training samples from the training samples containing the target plot categories, to obtain the training samples used for the batch training;

[0035] According to the training samples used for the batch training, a summary generation model is trained; the trained summary generation model is used to output a corresponding plot summary of the to-be-processed plot text based on the input to-be-processed plot text.

[0036] The drama text processing method, device, computer device, storage medium and computer program product can support generation of a drama summary, can reduce the case that the drama summary output produces an incorrect understanding of the drama text, and on this basis, by obtaining the number of classification manners used to determine the drama categories to which a plurality of training samples belong, and obtaining the number of drama categories configured for the training samples used in a batch of training, the number of classification manners and the number of drama categories can be used to determine the number of drama categories extracted from the drama categories of the plurality of training samples under each classification manner, and then for each classification manner, the target drama category can be extracted from the drama categories of the plurality of training samples under the classification manner according to the number of drama categories extracted, and the target training sample can be extracted from the training sample containing the target drama category, to obtain the training samples used in a batch of training, and the summary generation model is trained according to the training samples used in a batch of training. In the whole process, by the manner of first extracting the target drama category and then extracting the target training sample based on the target drama category, the prediction tasks of two levels of drama category prediction and drama summary prediction can be balanced, so that the model can learn the drama summary from the balanced prediction tasks of the two levels, can reduce the influence of biased data, can improve the generation effect of the drama summary under different drama categories, can obtain the summary generation model capable of outputting an accurate drama summary, and by inputting the drama text to be processed into the trained summary generation model, an accurate drama summary can be generated. BRIEF DESCRIPTION OF DRAWINGS

[0037] Figure 1 An application environment diagram of the drama text processing method in one embodiment;

[0038] Figure 2 A flowchart of the drama text processing method in one embodiment;

[0039] Figure 3 A flowchart of a character relationship to which a drama text belongs in one embodiment;

[0040] Figure 4 A flowchart of obtaining the training samples used in a batch of training in one embodiment;

[0041] Figure 5 A flowchart of obtaining the training samples used in a batch of training in another embodiment;

[0042] Figure 6 A schematic diagram of the network structure of the summary prediction network in one embodiment;

[0043] Figure 7 a schematic diagram of a model structure of an initial summary generation model in one embodiment;

[0044] Figure 8 a schematic diagram of a model structure of an initial summary generation model in another embodiment;

[0045] Figure 9 a schematic diagram of a model structure of an initial summary generation model in yet another embodiment;

[0046] Figure 10 a schematic diagram of a cascaded hierarchical information learning process when training a model in one embodiment;

[0047] Figure 11 a schematic diagram of a cascaded hierarchical information learning process when training a model in another embodiment;

[0048] Figure 12 a schematic diagram of a cascaded hierarchical information learning process when training a model in yet another embodiment;

[0049] Figure 13 a schematic diagram of a model structure of an initial summary generation model in yet another embodiment;

[0050] Figure 14 a scenario diagram of an application of a plot text processing method in one embodiment;

[0051] Figure 15 a structural block diagram of a plot text processing apparatus in one embodiment;

[0052] Figure 16 an internal structure diagram of a computer device in one embodiment. DETAILED DESCRIPTION

[0053] In order to make the purposes, technical solutions and advantages of the present application clearer, the present application is further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and should not be used to limit the present application.

[0054] The plot text processing method provided by the embodiments of the present application can be applied to, for example, Figure 1The application environment shown. Among them, the terminal 102 communicates with the server 104 through the network. The data storage system can store the data required by the server 104 to process. The data storage system can be set up separately, can be integrated on the server 104, or can be placed on the cloud or other servers. The server 104 obtains a plurality of training samples, the training samples including plot text, plot abstract, and at least one plot category to which the plot text belongs, each plot category being obtained by classifying the plot text according to a classification manner, obtaining the number of classification manners used to determine the plot categories to which the plurality of training samples belong, and obtaining the number of plot categories configured for a batch of training samples used for training. According to the number of classification manners and the number of plot categories, determine the plot category extraction number of plot categories extracted from each classification manner of the plot categories of the plurality of training samples, for each classification manner, according to the plot category extraction number, from the plot categories of the plurality of training samples under the classification manner, extract the target plot category, and from the training samples containing the target plot category, extract the target training sample, obtain the training sample used for a batch of training, and train the abstract generation model according to the training sample used for a batch of training. The trained abstract generation model is used to output the corresponding plot abstract of the plot text based on the input plot text. When receiving the plot abstract generation request of the terminal 102, the server 104 processes the plot text carried in the plot abstract generation request through the abstract generation model, generates the plot abstract and feeds back to the terminal 102.

[0055] Among them, the terminal 102 can be but not limited to various desktop computers, notebook computers, smart phones, tablet computers, Internet of Things devices and portable wearable devices, Internet of Things devices can be smart speakers, smart televisions, smart air conditioners, smart vehicle devices, etc. Portable wearable devices can be smart watches, smart bracelets, head-mounted devices, etc. The server 104 can be implemented by an independent server or a server cluster composed of multiple servers, or can be a cloud server or a node on a blockchain.

[0056] In one embodiment, as Figure 2 shown, a plot text processing method is provided, which can be executed by a terminal or a server alone, or by a terminal and a server cooperatively. In the embodiments of the present application, the method is applied to the server as an example, which includes the following steps:

[0057] Step 202, obtaining a plurality of training samples, the training samples including plot text, plot abstract, and at least one plot category to which the plot text belongs, each plot category being obtained by classifying the plot text according to a classification manner.

[0058] The training sample is a sample used for training the summary generation model. In the training sample, the plot text, the plot summary, and at least one plot category to which the plot text belongs are included. Each plot category is obtained by classifying the plot text according to a classification manner. The plot text refers to text used to describe the development of a plot in a movie, a TV series, a play, a novel, or other narrative works. For example, the plot text can specifically refer to a script used to describe the development of a plot in a movie or a TV series. The script refers to a document arranged by a scriptwriter for easy performance, mainly composed of time, place, characters, and dialogues between characters, and used for subsequent shooting guidance.

[0059] The plot summary refers to text information extracted from the plot text in terms of event description, which can summarize the plot text and generally includes time, place, characters, causes, processes, and results. The plot category refers to a category to which the plot text belongs, which is obtained by classifying the plot text according to a classification manner and can be one of a plurality of preset plot categories under the classification manner. The classification manner refers to a manner used to distinguish different plot texts, which can be configured according to actual application scenarios. For example, the classification manner can specifically be classification according to plot types, time backgrounds, plot occurrence background places, character relationships, etc.

[0060] The plot is one of the elements of the content of the plot text, which refers to a series of development processes of life events representing the relationships between characters in the plot text, and is a series of specific events showing the characters' personalities and representing the relationships between characters and the environment. In this embodiment, a plurality of preset plot types can be configured according to actual application scenarios. For example, when the classification manner is classification according to plot types, the plurality of preset plot categories under the classification manner can specifically be preset plot types such as a palace power struggle plot, a historical romance plot, a historical martial arts plot, a historical case plot, a historical life plot, a modern entrepreneurship plot, an office plot, etc.

[0061] The time background refers to a historical situation or a real environment that affects characters and events in the plot text. In this embodiment, a plurality of preset time backgrounds can be configured according to actual application scenarios. For example, when the classification manner is classification according to time backgrounds, the plurality of preset plot categories under the classification manner can specifically be preset time backgrounds such as a Republican era, a modern era, and an ancient era.

[0062] The character relationship refers to a relationship between at least two key characters appearing in the plot text. The key character refers to a main character that pushes the plot development appearing in the plot text. It can be understood that the key character can specifically be a character appearing most frequently in the plot text. In this embodiment, a plurality of preset character relationships can be configured according to actual application scenarios. For example, when the classification manner is classification according to character relationships, the plurality of preset plot categories in the classification manner can specifically be a father-son relationship, a father-daughter relationship, a mother-son relationship, a mother-daughter relationship, a husband-wife relationship, a couple relationship, a teacher-student relationship, a classmate relationship, and a colleague relationship.

[0063] Specifically, when the summary generation model needs to be trained, the server obtains a plurality of training samples, the training sample including a plot text, a plot summary, and at least one plot category to which the plot text belongs. Each plot category is obtained by classifying the plot text according to a classification manner.

[0064] In a specific application, the plot summary is obtained by pre-extracting a plot from the plot text, and the plot category is obtained by pre-classifying the plot text according to a classification manner, that is, when generating the training sample, the server first obtains the plot text, then extracts the plot summary from the plot text, and classifies the plot text according to at least one classification manner to obtain at least one plot category to which the plot text belongs.

[0065] In a specific application, when the plot text is obtained, the server can extract the plot summary by using a pre-trained natural language model. Specifically, the server can obtain the plot summary by asking the pre-trained natural language model. For example, the question sentence can be “The following is a plot text of a scene in a drama, please extract the plot summary of the scene, and do not make guesses or abstract descriptions, the plot text is XXX.” Wherein, XXX represents a specific plot text.

[0066] It can be understood that since the plot summary obtained at this time may have detail errors or character errors, it is usually necessary to further manually correct it, and after correction, the plot summary corresponding to the plot text can be obtained. The pre-trained natural language model can be configured according to actual application scenarios. For example, the pre-trained natural language model can specifically be a generative pre-training Transformer model, or other models based on the Transformer model.

[0067] In one specific application, after obtaining the plot text and determining the multiple preset plot categories under the classification manner, the server can also use the pre-trained natural language model to classify the plot text to obtain the plot category to which the plot text belongs. Specifically, the server can obtain the plot category to which the plot text belongs by asking the pre-trained natural language model. It can be understood that the plot category at this time is one of the multiple preset plot categories under the determined classification manner.

[0068] For example, taking classification by plot type as an example, the question sentence can be "The following is a plot text of a scene of a drama, according to the plot text, answer why the plot type of the plot text is, where the plot type includes plot type 1, plot type 2, plot type 3, plot type 4, and plot type 5, only answer the plot type that you are sure of, and answer that the plot type is not clear when the plot type is not clear. The plot text is: XXX". Wherein, XXX represents the specific plot text, and plot type 1, plot type 2, plot type 3, and plot type 4 are preset plot types, which can be configured according to actual application scenarios.

[0069] For example, taking classification by plot type as an example, the question sentence can be "The following is a plot text of a scene of a drama, according to the plot text, answer why the plot type of the plot text is, where the plot type includes plot type 1, plot type 2, plot type 3, plot type 4, and plot type 5, only answer the plot type that you are sure of, and answer that the plot type is not clear when the plot type is not clear. The plot text is: XXX". Wherein, XXX represents the specific plot text, and plot type 1, plot type 2, plot type 3, and plot type 4 are preset plot types, which can be configured according to actual application scenarios.

[0070] In one specific application, if the classification manner is classification by character relationship, the server needs to first determine at least two key characters whose relationship needs to be determined from the plot text, and then use the pre-trained natural language model to extract the character relationship. For example, taking at least two key characters whose relationship needs to be determined as key character A and key character B as an example, the question sentence can be "The following is a plot text of a scene of a drama, according to the plot text, answer what is the relationship between character A and character B, where the character relationship includes character relationship 1, character relationship 2, character relationship 3, character relationship 4, and character relationship 5, only answer the character relationship that you are sure of, and answer that the character relationship is not clear when the character relationship is not clear. The plot text is: XXX". Wherein, XXX represents the specific plot text, and character relationship 1, character relationship 2, character relationship 3, and character relationship 4 are preset character relationships, which can be configured according to actual application scenarios.

[0071] In one specific application, if the classification manner is to classify according to the character relationship, the flowchart for determining the character relationship to which the plot text belongs can be as shown in Figure 3 The specific steps include the following steps:

[0072] Step 302, the appearance frequency of the plot characters appearing in the plot text is counted to determine a plurality of plot characters and the appearance frequency of each plot character in the plurality of plot characters;

[0073] Step 304, at least two key characters are selected from the plurality of plot characters according to the appearance frequency of each plot character in the plurality of plot characters;

[0074] Step 306, each two key characters in the at least two key characters are taken as a group to obtain a plurality of key character pairs;

[0075] Step 308, for each key character pair, the relationship between the two key characters in the key character pair is predicted according to the plot text to obtain the key character relationship between the two key characters in the key character pair;

[0076] Step 310, when the number of the at least two key characters is two, the key character relationship between the two key characters is taken as the character relationship to which the plot text belongs;

[0077] Step 312, when the number of the at least two key characters is greater than two, the two key characters with the highest appearance frequency are determined according to the appearance frequency of each key character in the at least two key characters, and the key character relationship of the two key characters with the highest appearance frequency is taken as the character relationship to which the plot text belongs.

[0078] Step 204, the number of classification manners used to determine the plot categories to which a plurality of training samples belong is obtained, and the number of plot categories configured for a batch of training samples used for training is obtained.

[0079] The batch training refers to a training method in machine learning and deep learning, which involves updating the parameters of the model at the same time using a fixed number of training samples. The number of plot categories refers to the number of plot categories that need to be included in the training samples used for a batch of training. For example, if the number of plot categories is N, it means that the training samples used for a batch of training need to include N kinds of plot categories, where N is a positive integer. The number of plot categories can be configured according to the actual application scenario.

[0080] Specifically, after obtaining the plurality of training samples, the server obtains a number of classification manners used for determining the plot categories to which the plurality of training samples belong, and obtains a number of plot categories configured for a batch of training, so as to extract a target plot category according to the number of classification manners and the number of plot categories, and then extract target training samples from training samples containing the target plot category to obtain training samples used for the batch of training.

[0081] In step 206, the number of plot categories extracted from the plot categories under each classification manner of the plurality of training samples is determined according to the number of classification manners and the number of plot categories.

[0082] In the embodiment, the number of plot categories extracted from the plot categories under each classification manner of the plurality of training samples is determined according to the number of classification manners and the number of plot categories. In a specific application, when the number of plot categories is greater than or equal to the number of classification manners, the server calculates the ratio of the number of plot categories to the number of classification manners, and takes the ratio as the number of plot categories extracted from the plot categories under each classification manner of the plurality of training samples. When the number of plot categories is less than the number of classification manners, the server can directly take the number of plot categories as the number of plot categories extracted from the plot categories under each classification manner of the plurality of training samples.

[0083] In the embodiment, the number of plot categories extracted from the plot categories under each classification manner of the plurality of training samples is determined according to the number of classification manners and the number of plot categories. In a specific application, when the number of plot categories is greater than or equal to the number of classification manners, the server calculates the ratio of the number of plot categories to the number of classification manners, and takes the ratio as the number of plot categories extracted from the plot categories under each classification manner of the plurality of training samples. When the number of plot categories is less than the number of classification manners, the server can directly take the number of plot categories as the number of plot categories extracted from the plot categories under each classification manner of the plurality of training samples.

[0084] In step 208, for each classification manner, the target plot category is extracted from the plot categories under the classification manner of the plurality of training samples according to the number of plot categories extracted, and the target training sample is extracted from the training sample containing the target plot category to obtain the training sample used for the batch of training.

[0085] In the embodiment, for each classification manner, the server randomly extracts the number of plot categories extracted from the plot categories under the classification manner of the plurality of training samples as the target plot category according to the number of plot categories extracted, and extracts the target training sample from the training sample containing the target plot category to obtain the training sample used for the batch of training.

[0086] In a specific application, taking the number of classification manners as 2 and the number of plot categories greater than 2 as N as an example, the number of plot categories extracted can be N / 2. For each classification manner, the server randomly extracts N / 2 plot categories from the plot categories under the classification manner of the plurality of training samples as the target plot category.

[0087] At step 210, the summary generation model is trained according to the training samples used in a batch of training; the trained summary generation model is used to output a corresponding plot summary of the to-be-processed plot text based on the input to-be-processed plot text.

[0088] Specifically, after obtaining the training samples used in a batch of training, the server trains the summary generation model according to the training samples used in a batch of training; the trained summary generation model is used to output a corresponding plot summary of the to-be-processed plot text based on the input to-be-processed plot text.

[0089] In a specific application, for each target training sample in the training samples used in a batch of training, the server inputs the plot text in the target training sample into the initial summary generation model to perform plot summary prediction and plot category prediction, and performs prediction loss calculation according to the prediction result, the plot summary in the target training sample, and at least one plot category to which the plot text belongs, to obtain a prediction loss value corresponding to the target training sample; and then the parameters of the initial summary generation model are adjusted according to the prediction loss value corresponding to each target training sample in the training samples used in a batch of training, to obtain the summary generation model.

[0090] The plot text processing method described above, when training the summary generation model, introduces at least one plot category to which the plot text belongs in the training samples, combines the plot category into the summary generation model for learning, supports the generation of plot summaries, and can reduce the situation that the plot summary output produces an incorrect understanding of the plot text. On this basis, by obtaining the number of classification manners used to determine the plot categories to which a plurality of training samples belong, and obtaining the number of plot categories configured for the training samples used in a batch of training, the number of classification manners and the number of plot categories can be used to determine the plot category extraction number of plot categories extracted from the plot categories of a plurality of training samples under each classification manner, and then for each classification manner, the target plot category can be extracted from the plot categories of a plurality of training samples under the classification manner according to the plot category extraction number, and the target training sample can be extracted from the training samples containing the target plot category, to obtain the training samples used in a batch of training. The summary generation model is trained according to the training samples used in a batch of training. In the entire process, by first extracting the target plot category and then extracting the target training sample based on the target plot category, the balance of the two levels of prediction tasks of plot category prediction and plot summary prediction can be achieved, so that the model can jointly learn the plot summary from the balance of the two levels of prediction tasks, can reduce the influence of biased data, can improve the generation effect of the plot summary under different plot categories, can obtain a summary generation model that can output an accurate plot summary, and can generate an accurate plot summary by inputting the to-be-processed plot text into the trained summary generation model.

[0091] In an embodiment, the number of extracted plot categories from the plot categories under each classification mode of the plurality of training samples is determined according to the number of classification modes and the number of plot categories, including:

[0092] determining a ratio between the number of plot categories and the number of classification modes;

[0093] taking the ratio as the number of extracted plot categories from the plot categories under each classification mode of the plurality of training samples.

[0094] Specifically, when the number of plot categories is greater than or equal to the number of classification modes, the server calculates the ratio between the number of plot categories and the number of classification modes, and takes the ratio as the number of extracted plot categories from the plot categories under each classification mode of the plurality of training samples. For example, when the number of classification modes is 2 and the number of plot categories is N, which is greater than 2, the number of extracted plot categories can be N / 2.

[0095] It should be noted that the number of classification modes can also be 1. When the number of classification modes is 1, the ratio between the number of plot categories and the number of classification modes is the same as the number of plot categories, that is, the number of plot categories is directly taken as the number of extracted plot categories from the plot categories under each classification mode of the plurality of training samples.

[0096] In this embodiment, by first determining the ratio between the number of plot categories and the number of classification modes, the ratio can be used to determine the number of extracted plot categories. By taking the ratio as the number of extracted plot categories, the number of extracted plot categories under each classification mode can be balanced, thereby reducing the influence of biased data during model training and improving the generation effect of plot summaries under different plot categories.

[0097] In an embodiment, for each classification mode, a target plot category is extracted from the plot categories of the plurality of training samples under the classification mode according to the number of extracted plot categories, and a target training sample is extracted from the training samples containing the target plot category, to obtain the training samples used for a batch of training, including:

[0098] For each classification mode, a target plot category is extracted from the plot categories of the plurality of training samples under the classification mode according to the number of extracted plot categories, and a training sample number configured for the target plot category is obtained;

[0099] According to the training sample number, a target training sample is extracted from the training samples containing the target plot category, to obtain the training samples used for a batch of training.

[0100] The training sample quantity is the quantity of training samples that need to be extracted from the training samples containing the target plot category. Specifically, the training sample quantity can be a positive integer greater than or equal to 1.

[0101] Specifically, for each classification mode, the server will extract a plot category extraction quantity of plot categories from the plot categories of the plurality of training samples under the classification mode, as the target plot category, obtain the training sample quantity configured for the target plot category, extract the target training sample from the training samples containing the target plot category according to the training sample quantity, and obtain the training sample used for a batch of training.

[0102] In a specific application, as shown in FIG. 8, when the quantity of target plot categories is one, the server can directly extract the target training sample from the training samples containing the target plot category according to the training sample quantity configured for the target plot category. Figure 4 As shown in FIG. 8, when the quantity of target plot categories is multiple, for each target plot category, the server needs to obtain the training sample quantity configured for the target plot category, extract the target training sample from the training samples containing the target plot category according to the training sample quantity, and obtain the training sample used for a batch of training. It can be understood that when the quantity of target plot categories is multiple, for each target plot category in the multiple target plot categories, the server can extract the corresponding target training sample according to the target plot category, and obtain the training sample used for a batch of training by collecting the target training samples corresponding to the multiple target plot categories.

[0103] In a specific application, when the quantity of target plot categories is multiple, the training sample quantity configured for each target plot category can be the same or different, which can be configured according to the actual application scenario.

[0104] In a specific application, when the training sample quantity is 1, the server directly extracts one training sample from the training samples containing the target plot category as the target training sample. When the training sample quantity is greater than 1, the server needs to obtain the summary category to which the plot summary in each training sample containing the target plot category belongs, and then extract the target training sample from the training samples containing the target plot category according to the training sample quantity and the quantity of summary categories to which the plot summaries belong, to obtain the training sample used for a batch of training.

[0105] In the embodiment, for each classification manner, the target plot category is extracted from the plot categories of the plurality of training samples under the classification manner according to the plot category extraction quantity, so that the target plot category under each classification manner is balanced, and then the target training sample is extracted from the training samples containing the target plot category by using the obtained training sample quantity configured for the target plot category, so that the training sample used for a batch of training is obtained. By extracting the target plot category first and then extracting the target training sample based on the target plot category, the balance of the prediction tasks of the two levels of plot category prediction and plot summary prediction can be achieved.

[0106] In one embodiment, extracting the target training sample from the training samples containing the target plot category according to the training sample quantity to obtain the training sample used for a batch of training comprises:

[0107] Clustering the plot summaries in each of the training samples containing the target plot category to obtain the summary category to which the plot summary in each of the training samples containing the target plot category belongs;

[0108] According to the training sample quantity and the number of summary categories to which the plot summaries belong, determining a summary category extraction quantity of extracting a summary category from the summary categories to which the plot summaries belong;

[0109] Extracting a target summary category from the summary categories to which the plot summaries belong according to the summary category extraction quantity, and extracting a target training sample from the training sample containing the target plot category and the plot summary belonging to the target summary category to obtain the training sample used for a batch of training.

[0110] Wherein, clustering is an unsupervised learning method for grouping samples in a data set so that the similarity of samples in the same group is high and the similarity of samples between different groups is low. In the embodiment, the plot summaries in each of the training samples containing the target plot category are grouped so that the similarity of plot summaries in the same group is high and the similarity of plot summaries between different groups is low. The plot summaries in the same group belong to the same summary category, and the plot summaries in different groups belong to different summary categories.

[0111] The summary category refers to a category to which a plot summary is attributed when the plot summary in each training sample containing the target plot category is classified by clustering. It should be noted that the summary category is determined by clustering and can be understood as an abstract category to which the plot summary is attributed, and the purpose is only to group the plot summaries in each training sample containing the target plot category. For example, if the plot summaries in each training sample containing the target plot category are divided into N groups by clustering, each group can be understood as a summary category. The summary category extraction number refers to the number of summary categories that need to be extracted from the summary categories to which the plot summaries are attributed.

[0112] Specifically, the server can obtain the summary categories to which the plot summaries in each training sample containing the target plot category are attributed by clustering the plot summaries in each training sample containing the target plot category. On this basis, by comparing the number of training samples and the number of summary categories to which the plot summaries are attributed, the number of summary categories to be extracted from the summary categories to which the plot summaries are attributed can be determined, and then the number of summary categories to be extracted from the summary categories to which the plot summaries are attributed can be randomly extracted as target summary categories according to the number of summary categories to be extracted, and the target training sample can be extracted from the training sample containing the target plot category and the plot summary attributed to the target summary category, to obtain the training sample used for a batch of training.

[0113] In a specific application, when the number of training samples is less than or equal to the number of summary categories to which the plot summaries are attributed, the flowchart for obtaining the training sample used for a batch of training can be as shown in Figure 5 , and specifically includes the following steps:

[0114] Step 502, taking the number of training samples as the number of summary categories to be extracted from the summary categories to which the plot summaries are attributed;

[0115] Step 504, randomly extracting the number of summary categories to be extracted from the summary categories to which the plot summaries are attributed as target summary categories according to the number of summary categories to be extracted;

[0116] Step 506, extracting a training sample containing the target plot category and the plot summary attributed to the target summary category as a target training sample to obtain the training sample used for a batch of training.

[0117] For example, when the number of training samples is 1 and the number of summary categories to which the plot summaries are attributed is a positive integer N greater than 1, the server can directly determine that the number of summary categories to be extracted is 1.

[0118] In a specific application, when the number of training samples is greater than the number of summary categories to which the plot summaries belong, a flowchart for obtaining the training samples used by the batch training can be as shown in Figure 5 , and specifically includes the following steps:

[0119] Step 508, the number of summary categories to which the plot summaries belong is taken as the number of summary categories extracted from the summary categories to which the plot summaries belong;

[0120] Step 510, the number ratio between the number of training samples and the number of summary categories to which the plot summaries belong is calculated;

[0121] Step 512, according to the number ratio, n training samples are extracted from the training samples containing the target plot category and belonging to the target summary category as target training samples, so that the number of extracted training samples meets the demand.

[0122] For example, when the number of training samples is 2n, and the number of summary categories to which the plot summaries belong is n, n is a positive integer greater than or equal to 1, the server will take n as the number of summary categories extracted, and calculate the number ratio between the number of training samples and the number of summary categories to which the plot summaries belong (2n / n=2), and extract 2 training samples from the training samples containing the target plot category and belonging to the target summary category as target training samples.

[0123] In this embodiment, by first clustering the plot summaries of each training sample to determine the summary categories to which the plot summaries of each training sample belong, then using the number of training samples and the number of summary categories to which the plot summaries belong to determine the number of summary categories extracted, and finally using the number of summary categories extracted to extract the target summary category from the summary categories to which the plot summaries belong, and extracting the target training sample from the training samples containing the target plot category and belonging to the target summary category, the training samples used by the batch training are obtained, which can further balance the plot summaries under the condition of balanced plot categories, and can balance the prediction tasks of the two levels of plot category prediction and plot summary prediction, so that the model can jointly learn the plot summaries from the balanced prediction tasks of the two levels, can reduce the influence of biased data, and can improve the generation effect of plot summaries under different plot categories.

[0124] In one embodiment, clustering the plot summaries in each training sample containing the target plot category to obtain the summary categories to which the plot summaries in each training sample containing the target plot category belong includes:

[0125] performing representation extraction on the plot summary in the training sample to obtain a representation of the plot summary in the training sample;

[0126] performing clustering on the representation of the plot summary of each training sample containing the target plot category to obtain a summary category to which the plot summary in each training sample containing the target plot category belongs.

[0127] In this embodiment, useful information is extracted from the plot summary in the training sample and converted into a format or a set of features. That is, the representation of the plot summary in the training sample in this embodiment can be a feature vector, etc.

[0128] Specifically, for each training sample containing the target plot category, the server will perform representation extraction on the plot summary in the training sample to obtain a representation of the plot summary in the training sample, and further perform clustering on the representation of the plot summary of each training sample containing the target plot category to obtain a summary category to which the plot summary in each training sample containing the target plot category belongs. Specifically, plot summaries with high representation similarity are divided into a group, and plot summaries with low representation similarity are divided into different groups.

[0129] In specific applications, the representation extraction method can be to encode the data into a format that the model can understand (such as using one-hot encoding or word embedding to encode the plot summary), or to use a pre-trained natural language model to convert the plot summary into a numerical vector form. The pre-trained natural language model can be trained according to the actual application scenario. For example, the pre-trained natural language model can be a pre-trained Text-to-Vector model, which can convert Chinese text into a numerical vector form and capture semantic information in the text.

[0130] In specific applications, the server can use a pre-trained clustering algorithm to cluster the representation of the plot summary of each training sample containing the target plot category to obtain a summary category to which the plot summary in each training sample containing the target plot category belongs. The pre-trained clustering algorithm can be configured according to the actual application scenario. Specifically, it can be a K-Means clustering, hierarchical clustering, density clustering, model-based clustering, etc.

[0131] Among them, K-Means clustering is a simple and widely used clustering algorithm, which selects cluster centers and assigns samples to the nearest cluster center by iteration. Hierarchical clustering constructs a hierarchically nested cluster tree by gradually merging or splitting samples. Density clustering, such as DBSCAN (Density-Based Spatial Clustering of Applications with Noise), clusters based on the connectivity of sample density, suitable for discovering clusters of arbitrary shape. Model-based clustering, such as Gaussian mixture model, assumes that data is mixed by multiple probability distributions, and clusters by estimating the parameters of these distributions.

[0132] In this embodiment, for each training sample containing the target plot category, the plot summary in the training sample is extracted by the plot summary representation, and then the plot summary in each training sample containing the target plot category is determined by clustering the plot summary representation of each training sample containing the target plot category, so as to realize the extraction of the plot summary according to the summary category, and realize the further balance of the plot summary in the case of balanced plot categories.

[0133] In one embodiment, the training of the summary generation model according to the training samples used in a batch of training includes:

[0134] For each target training sample in the training samples used in a batch of training, input the plot text in the target training sample into the initial summary generation model to predict the plot summary and the plot category, obtain the predicted summary and the category features of at least one predicted category, and each predicted category is obtained by classifying the plot text according to a classification method.

[0135] According to the predicted summary, the category features of at least one predicted category, the plot summary in the target training sample, and at least one predicted plot category, the prediction loss is calculated to obtain the prediction loss value corresponding to the target training sample.

[0136] According to the prediction loss value corresponding to each target training sample in the training samples used in a batch of training, the parameters of the initial summary generation model are adjusted to obtain the summary generation model.

[0137] The predicted summary refers to a predicted summary output by the initial summary generation model based on the plot text in the target training sample. The category feature of the predicted category refers to a feature representing the predicted category. For example, the category feature of the predicted category can specifically refer to a vector representing the predicted category. It can be understood that the category feature of the predicted category output by the initial summary generation model can uniquely determine the predicted category. The predicted category refers to a predicted plot category that can be uniquely determined based on the category feature of the predicted category output by the initial summary generation model. It can be understood that each predicted category is obtained by classifying the plot text according to a classification manner, and each plot category is obtained by classifying the plot text according to a classification manner.

[0138] Specifically, for each target training sample in a batch of training training samples, the server inputs the plot text in the target training sample into the initial summary generation model to predict the plot summary and the plot category, and obtains the predicted summary and the category feature of at least one predicted category. In a specific application, after the plot text in the target training sample is input into the initial summary generation model, the initial summary generation model encodes the plot text in the target training sample to obtain an encoding vector corresponding to the plot text, and then uses the encoding vector corresponding to the plot text to predict the plot summary and the plot category, and obtain the predicted summary and the category feature of at least one predicted category.

[0139] In a specific application, the initial summary generation model can encode each word in the plot text in the target training sample to obtain a word encoding corresponding to each word in the plot text in the target training sample, and the encoding vector corresponding to the plot text in the target training sample includes the word encoding corresponding to each word in the plot text in the target training sample.

[0140] For example, the initial summary generation model can encode each word in the plot text in the target training sample by using a preset dictionary. The preset dictionary can be configured according to an actual application scenario. Each Chinese character or letter in the preset dictionary has a unique word encoding. Therefore, by using the preset dictionary to encode each word in the plot text in the target training sample, each word in the plot text in the target training sample is actually mapped to a word encoding in the preset dictionary. For further example, the preset dictionary can be a dictionary set based on one-hot encoding. Each Chinese character or letter in the preset dictionary corresponds to a word encoding with a length of 1*Nword, where Nword is the total number of Chinese characters and letters in the preset dictionary.

[0141] In one specific application, when predicting the plot summary by using the encoding vector corresponding to the plot text in the target training sample, the server can generate the plot summary by word-by-word prediction, i.e., the first word in the predicted summary is directly predicted according to the encoding vector corresponding to the plot text, and for each word after the first word in the predicted summary, the server can combine the encoding vector corresponding to the plot text and the last predicted word to predict the word in the predicted summary. In this way, the last prediction result can be combined to predict at each time, and the generation of the predicted summary can be realized on the basis of the previous text.

[0142] In one specific application, taking the number of words in the predicted summary as N words or less and the same word prediction network used for each word prediction as an example, the network structure of the summary prediction network in the initial summary generation model can be as shown in Figure 6 , which includes an encoding network, N word prediction networks, and a fully connected layer. The encoding network is used to encode the plot text in the target training sample to obtain the encoding vector corresponding to the plot text in the target training sample. The first word prediction network is used to combine the fully connected layer and perform word prediction based on the encoding vector corresponding to the plot text in the target training sample, and output the prediction result of the first word in the predicted summary (as shown in Figure 6 , the predicted results are represented by id, and specifically can be the prediction probability of the first word). From the second word prediction network, it is used to combine the fully connected layer and perform prediction based on the encoding vector corresponding to the plot text in the target training sample and the last predicted result, and output the prediction result of the word at the corresponding position in the predicted summary, which can be the prediction probability of the word at the corresponding position. Each word prediction network can output the probability of the word at the corresponding position in the predicted summary belonging to each word in the preset dictionary.

[0143] The word prediction network can be configured according to the actual application scenario, for example, the word prediction network can be a Transformer model, and the network structure of the summary prediction network in the initial summary generation model in this embodiment is mainly a structure of multiple stacked Transformer models. Further, the Transformer model can be a DecoderLayer (decoding layer) in the Transformer model, mainly including a normalization layer, an attention layer, and a multi-layer perceptron, etc.

[0144] In one specific application, when predicting the plot category by using the encoding vector corresponding to the plot text in the training sample for the target, in order to assist the prediction of the plot category to the prediction of the plot summary, the intermediate feature output by the initial summary generation model in the prediction of the plot summary part can be used to predict the plot category. For example, the feature output by the first word prediction network in the summary prediction network of the initial summary generation model can be used to predict the plot category, and the model structure of the initial summary generation model can be as shown in Figure 7 , the output of the first word prediction network is input into the plot classification network (only one plot classification network corresponding to one classification mode is shown in Figure 7 ), and the plot classification network can obtain the category feature of the predicted category by processing the output of the first word prediction network.

[0145] Further, in the case of multiple classification modes, each classification mode corresponds to a plot classification network, and the model structure of the initial summary generation model can be as shown in Figure 8 , the output of the first word prediction network is input into each plot classification network, and for each plot classification network, the plot classification network can obtain the category feature of the predicted category by processing the output of the first word prediction network.

[0146] In one specific application, the plot classification network can obtain the category feature of the predicted category by classifying based on the output of the first word prediction network, and the network structure of the plot classification network can be a pooling layer and a fusion layer, i.e., the output of the first word prediction network is first pooled by the pooling layer, and then the pooled feature is fused by the fusion layer, so as to obtain the category feature of the predicted category. In this embodiment, the pooling mode of the pooling layer is not limited, which can be average pooling, maximum pooling, etc.

[0147] In one specific application, in the case where the network structure of the plot classification network is a pooling layer and a fusion layer, the initial summary generation model can include a summary prediction network and a plot classification network, as shown in Figure 9 , (only one plot classification network corresponding to one classification mode is shown in Figure 9 ), wherein the summary prediction network includes an encoding network, a plurality of word prediction networks, and a full connection layer, the plot classification network includes a pooling layer and a fusion layer, and the output of the first word prediction network is input into the plot classification network.

[0148] Specifically, after obtaining the predicted summary and the category feature of the at least one predicted category, the server performs prediction loss calculation according to the predicted summary, the category feature of the at least one predicted category, the plot summary in the target training sample and the at least one plot category to which the plot text belongs, to obtain a prediction loss value corresponding to the target training sample, and adjusts the parameters of the initial summary generation model according to the prediction loss value corresponding to each target training sample in a batch of training samples, to obtain the summary generation model.

[0149] In this embodiment, by inputting the plot text in the target training sample into the initial summary generation model to perform plot summary prediction and plot category prediction for each target training sample, the predicted summary and the category feature of the at least one predicted category can be obtained, and then the plot summary, the predicted summary, the at least one plot category to which the plot text belongs and the category feature of the at least one predicted category can be used to perform prediction loss calculation to obtain a prediction loss value corresponding to the target training sample, so that the summary generation model can be obtained by adjusting the parameters of the initial summary generation model according to the prediction loss value corresponding to each target training sample in a batch of training samples. During the training of the initial summary generation model, the prediction of the plot category is cascaded into the initial summary generation model, so that the initial summary generation model can understand the correct information related to the summary from the plot category and support plot summary generation, and the situation that the plot summary output produces an incorrect understanding of the plot text can be avoided. Therefore, the summary generation model capable of outputting an accurate plot summary can be obtained, and the accurate plot summary can be generated by inputting the plot text to be processed into the trained summary generation model.

[0150] In one embodiment, the prediction loss calculation according to the predicted summary, the category feature of the at least one predicted category, the plot summary in the target training sample and the at least one plot category to which the plot text belongs, to obtain a prediction loss value corresponding to the target training sample, comprises:

[0151] calculating a summary prediction loss value according to the predicted summary and the plot summary in the target training sample, and calculating a category prediction loss value according to the category feature of the at least one predicted category and the at least one plot category to which the plot text belongs;

[0152] aggregating the summary prediction loss value and the category prediction loss value to obtain the prediction loss value corresponding to the target training sample.

[0153] The summary prediction loss value refers to the error between the plot summary and the predicted summary. The category prediction loss value refers to the error between the category feature of the at least one predicted category and the feature of the at least one plot category to which the plot text belongs.

[0154] Specifically, in calculating the prediction loss value, the server calculates the summary prediction loss value according to the prediction summary and the plot summary in the target training sample, and for each classification manner, takes the plot category classified according to the classification manner and the prediction category predicted according to the classification manner as the loss calculation parameters of the classification manner, to calculate the prediction loss value of the classification manner corresponding to the classification manner, and then obtains the category prediction loss value according to the prediction loss value of each classification manner corresponding to the classification manner, and finally aggregates the summary prediction loss value and the category prediction loss value to obtain the prediction loss value corresponding to the target training sample.

[0155] In a specific application, the prediction summary output by the initial summary generation model can be a prediction summary vector, and in calculating the summary prediction loss value, the server can first encode the plot summary in the target training sample to obtain the encoding vector corresponding to the plot summary in the target training sample, and then determine the summary prediction loss value according to the encoding vector corresponding to the plot summary in the target training sample and the prediction summary vector.

[0156] In a specific application, in calculating the summary prediction loss value, the server can first encode each word in the plot summary in the target training sample through a preset dictionary to obtain the word encoding corresponding to each word in the plot summary in the target training sample, and determine the prediction probability corresponding to each word in the prediction summary, and then for each word in the plot summary in the target training sample, calculate the prediction loss value corresponding to the word according to the word encoding corresponding to the word and the prediction probability of the word at the same position in the prediction summary, and finally obtain the summary prediction loss value according to the prediction loss value corresponding to each word in the plot summary in the target training sample.

[0157] In a specific application, for each word in the plot summary in the target training sample, the server can determine whether the word at the same position in the prediction summary is the same as the word according to the word encoding corresponding to the word and the prediction probability of the word at the same position in the prediction summary, that is, the server can determine the classification loss corresponding to the word according to the prediction probability, so that the server can obtain the summary prediction loss value by combining the classification loss value corresponding to each word in the plot summary in the target training sample. In a specific application, the server can take the average of the classification loss values corresponding to each word in the plot summary in the target training sample as the summary prediction loss value.

[0158] In one specific application, the summary prediction loss value in the embodiment, which can be specifically cross-entropy loss value, that is, the classification loss value of each word in the predicted plot summary, needs to be explained that the probability of the class here comes from each word in the preset dictionary, and each word can be considered as a class. The formula for calculating the summary prediction loss value can be:

[0159] ;

[0160] Among them, represents the word encoding corresponding to each word in the plot summary, for each word in the plot summary, the word encoding of the word, is determined according to the position of the word in the preset dictionary, which can be specifically one-hot (one-hot encoding) label, when it is a certain word in the preset dictionary, the label of the position corresponding to the word in the preset dictionary in the word encoding is 1, and the labels of the other positions are 0, represents the predicted probability corresponding to each word in the predicted summary, that is, in the case of taking as the actual probability distribution of the plot summary, the summary prediction loss value is determined by comparing the predicted probability distribution of the predicted summary .

[0161] In a specific application, in the process of aggregating the summary prediction loss value and the category prediction loss value, the server can directly superimpose the summary prediction loss value and the category prediction loss value to obtain the prediction loss value corresponding to the target training sample, or can obtain the prediction loss value corresponding to the target training sample by pre-setting the weighting coefficients corresponding to the summary prediction loss value and the category prediction loss value, and then weighting and summing the summary prediction loss value and the category prediction loss value. The weighting coefficients can be configured according to the actual application scene.

[0162] It can be understood that because the plot summary prediction is the main task and the plot category prediction is the auxiliary task, the weighting coefficient corresponding to the summary prediction loss value is usually much larger than the weighting coefficient corresponding to the category prediction loss value. For example, the weighting coefficient corresponding to the summary prediction loss value can be set to 1, and the weighting coefficient corresponding to the category prediction loss value can be set to a (a positive number less than 1), and the prediction loss value corresponding to the target training sample Loss1= Loss_HH + a*Loss_HL, where Loss_HH represents the summary prediction loss value, and Loss_HL represents the category prediction loss value.

[0163] In the embodiment, the summary prediction loss value and the category prediction loss value are used to determine the prediction loss value corresponding to the target training sample by calculating the summary prediction loss value according to the predicted summary and the plot summary in the target training sample, and calculating the category prediction loss value according to the category features of at least one prediction category and at least one plot category to which the plot text belongs.

[0164] In one embodiment, calculating the category prediction loss value according to the category features of at least one prediction category and at least one plot category to which the plot text belongs includes:

[0165] For each classification manner, the plot category classified according to the classification manner and the category features of the prediction category predicted according to the classification manner are taken as the loss calculation parameters of the classification manner;

[0166] The prediction loss value corresponding to the classification manner is calculated according to the plot category and the category features of the prediction category in the loss calculation parameters;

[0167] The category prediction loss value is obtained according to the prediction loss value corresponding to each classification manner.

[0168] Specifically, since each plot category is obtained by classifying the plot text according to a classification manner, and each prediction category is obtained by predicting the plot text according to a classification manner, for each classification manner, the corresponding plot category and prediction category of the classification manner exist, therefore, for each classification manner, the server takes the plot category classified according to the classification manner and the category features of the prediction category predicted according to the classification manner as the loss calculation parameters of the classification manner, then extracts the features of the plot category in the loss calculation parameters to obtain the plot category features corresponding to the plot category in the loss calculation parameters, calculates the prediction loss value corresponding to the classification manner according to the plot category features and the category features of the prediction category in the loss calculation parameters, and finally obtains the category prediction loss value according to the prediction loss value corresponding to each classification manner.

[0169] In a specific application, taking at least one plot category as two plot categories (plot category 1 and plot category 2), where the plot category 1 is obtained by classifying the plot text according to the classification manner 1, and the plot category 2 is obtained by classifying the plot text according to the classification manner 2, and taking at least one prediction category as two prediction categories (prediction category 1 and prediction category 2), where the prediction category 1 is obtained by predicting the plot text according to the classification manner 1, and the prediction category 2 is obtained by predicting the plot text according to the classification manner 2, as an example, the calculation process of the category prediction loss value in the embodiment is illustrated.

[0170] In the calculation of the category prediction loss value, for the classification mode 1, the server takes the category feature of the plot category 1 classified according to the classification mode 1 and the prediction category 1 predicted according to the classification mode 1 as the loss calculation parameter of the classification mode 1, then extracts the feature of the plot category 1 to obtain the plot category feature corresponding to the plot category 1, and calculates the prediction loss value corresponding to the classification mode 1 according to the plot category feature and the category feature of the prediction category 1. For the classification mode 2, the server takes the category feature of the plot category 2 classified according to the classification mode 2 and the prediction category 2 predicted according to the classification mode 2 as the loss calculation parameter of the classification mode 2, then extracts the feature of the plot category 2 to obtain the plot category feature corresponding to the plot category 2, and calculates the prediction loss value corresponding to the classification mode 2 according to the plot category feature and the category feature of the prediction category 2.

[0171] In a specific application, after obtaining the prediction loss value corresponding to each classification mode, the server can directly take the prediction loss value corresponding to each classification mode as the category prediction loss value, or can calculate the average of the prediction loss value corresponding to each classification mode and take the calculated average as the category prediction loss value. In this embodiment, this is not limited, as long as the category prediction loss value is obtained based on the prediction loss value corresponding to each classification mode.

[0172] In this embodiment, for each classification mode, by taking the category feature of the plot category classified according to the classification mode and the prediction category predicted according to the classification mode as the loss calculation parameter of the classification mode, the category feature of the plot category and the prediction category in the loss calculation parameter can be used to calculate the prediction loss value corresponding to the classification mode, and then the category prediction loss value can be determined according to the prediction loss value corresponding to each classification mode.

[0173] In one embodiment, calculating the prediction loss value corresponding to the classification mode according to the category feature of the plot category and the prediction category in the loss calculation parameter comprises:

[0174] encoding the plot category in the loss calculation parameter to obtain the plot category feature corresponding to the plot category in the loss calculation parameter;

[0175] calculating the similarity loss value between the plot category feature and the category feature of the prediction category in the loss calculation parameter, and taking the similarity loss value as the prediction loss value corresponding to the classification mode.

[0176] The similarity loss value is used to measure the similarity between the plot category feature and the category feature of the predicted category. It can be understood that the more similar the plot category feature and the category feature of the predicted category, the smaller the similarity loss value, and the greater the gap between the plot category feature and the category feature of the predicted category, the greater the similarity loss value.

[0177] Specifically, when calculating the prediction loss value corresponding to the classification mode, the server will first encode the plot category in the loss calculation parameter to obtain the plot category feature corresponding to the plot category in the loss calculation parameter, and then calculate the similarity loss value between the plot category feature and the category feature of the predicted category in the loss calculation parameter. The similarity loss value is used as the prediction loss value corresponding to the classification mode. It can be understood that since the plot category feature represents the plot category and the category feature of the predicted category represents the predicted category, the similarity between the plot category and the predicted category can be determined by comparing the similarity between the plot category feature and the category feature of the predicted category. By loss calculation, the similarity between the predicted category and the plot category can be improved, which can make the initial summary generation model output contain more such plot categories, which is conducive to accurate plot summary prediction.

[0178] In a specific application, the loss calculation formula used in the calculation of the similarity loss value in the embodiment can be configured according to the actual application scene, as long as the similarity loss calculation can be realized. For example, the similarity loss in the embodiment can be Contrastive loss loss (contrastive loss), that is, the similarity between the plot category feature and the category feature of the predicted category in the loss calculation parameter is measured by Contrastive loss loss. In the embodiment, the input of Contrastive loss loss is a sample pair composed of the plot category feature and the category feature of the predicted category in the loss calculation parameter, and the label is whether the sample pair belongs to the same category, that is, whether the plot category and the category feature representing the predicted category in the loss calculation parameter are the same.

[0179] In one specific application, the formula for calculating the similarity loss value can be:

[0180] ;

[0181] wherein, is the abbreviation of the function , which represents the vector after the input is mapped, that is, the plot category feature in the embodiment, , which represents the vector after the input is mapped, that is, the category feature of the predicted category in the embodiment, , and all represent plot text, 1{·} is an indicator function that returns 1 when the input is true, and otherwise returns 0, that is, when the plot category in the loss calculation parameter and the predicted category represented by the category feature of the predicted category are the same, return 1, otherwise return 0, m is a pre-configured hyperparameter, which can be configured according to the actual application scenario, indicating that the distance between different class samples should exceed this value.

[0182] In this embodiment, by encoding the plot category, the plot category feature is obtained, and the similarity loss value between the two features can be calculated by using the plot category feature and the category feature of the predicted category, and then the prediction loss value corresponding to the classification mode can be determined by using the similarity loss value.

[0183] In one embodiment, according to the prediction loss value corresponding to each target training sample in the training sample used in a batch training, the parameters of the initial abstract generation model are adjusted to obtain an abstract generation model including:

[0184] According to the prediction loss value corresponding to each target training sample in the training sample used in a batch training, the parameters of the initial abstract generation model are adjusted to obtain a model after parameter adjustment;

[0185] According to the total number of samples of the plurality of training samples and the number of training samples used in a batch training, the number of batch trainings in a round of training is determined;

[0186] According to the number of batch trainings, the model after parameter adjustment is trained in multiple batches to obtain an abstract generation model.

[0187] In machine learning and deep learning, one round of training refers to the entire training data set being used for a complete training process. The number of batch trainings in a round of training refers to the number of batches required to complete a traversal of the entire training data set in machine learning and deep learning. In this embodiment, it refers to the number of batches required to complete a traversal of the plurality of training samples. Each batch uses a certain number of training samples, that is, the training samples used in a batch training, which are used to update the weights of the model at the same time. It can be understood that the batch size is an important hyperparameter that affects many aspects of model training.

[0188] Specifically, after obtaining the prediction loss value corresponding to each target training sample in the training samples used in a batch training, the server calculates the prediction loss value corresponding to the batch training according to the prediction loss value corresponding to each target training sample in the training samples used in the batch training, returns the prediction loss value corresponding to the batch training to the initial summary generation model, calculates the parameter gradient of each parameter of the initial summary generation model, updates each parameter according to the parameter gradient of each parameter to obtain a model with adjusted parameters, and determines the batch training times of a round of training according to the total number of training samples and the number of training samples used in a batch training, and performs multiple batch trainings on the model with adjusted parameters according to the batch training times to obtain the summary generation model.

[0189] In a specific application, after determining the batch training times of a round of training, the server can determine the number of batch trainings that still need to be performed according to the batch training times, i.e., the batch training times minus 1, and then can perform multiple batch trainings on the model with adjusted parameters according to the number of batch trainings that still need to be performed to obtain the summary generation model.

[0190] In a specific application, for each batch training, target training samples need to be re-extracted from the plurality of training samples as training samples used in the batch training for model training. It can be understood that in the process of extracting target training samples, the target plot categories and the target training samples are randomly extracted, and therefore, for each batch training, although the number of extracted plot categories is the same, the extracted target plot categories and target training samples can not be completely the same.

[0191] In a specific application, taking the total number of training samples as N and the number of training samples used in a batch training as bs as an example, the batch training times of a round of training is determined as M=N / bs according to the total number of training samples N and the number of training samples used in a batch training bs, and the number of batch trainings that still need to be performed is determined as M-1 by subtracting the batch training of this training, and then the summary generation model can be obtained by performing M-1 trainings on the model with adjusted parameters.

[0192] In one specific application, on the basis of completing one round of iterative training, the server can also continue to perform multiple rounds of iterative training with the multiple training samples as full samples, and train to obtain the summary generation model. In this regard, the full samples are processed once in each round of iterative training, and the iterative training is ended until a certain round of iteration meets the iteration stop condition, and the summary generation model is obtained. The iteration stop condition can be configured according to the actual application scenario. For example, the iteration stop condition can be that the prediction loss value obtained in the current round of iteration converges, that is, the loss value difference is less than the difference threshold compared with the prediction loss value obtained in the last iteration. The prediction loss value obtained in each round of iteration can be the average of the prediction loss values corresponding to each batch of training in the current round of iteration.

[0193] In this embodiment, by adjusting the parameters of the initial summary generation model according to the prediction loss value corresponding to each target training sample in the training sample used in each batch of training, a model with adjusted parameters can be obtained, and then based on the determination of the batch training times of one round of training, the model with adjusted parameters can be trained in multiple batches using the batch training times to obtain the summary generation model, and the accuracy of the summary generation model can be improved by means of multiple batch training.

[0194] In one embodiment, the plot text processing method further comprises:

[0195] Obtaining the plot text to be processed;

[0196] Inputting the plot text to be processed into the trained summary generation model to obtain the plot summary corresponding to the plot text to be processed.

[0197] Specifically, in the case of obtaining the summary generation model, when it is necessary to process the plot text, the server can obtain the plot text to be processed and input the plot text to be processed into the trained summary generation model, so as to obtain the plot summary corresponding to the plot text to be processed.

[0198] In this embodiment, since the summary generation model capable of outputting accurate plot summaries is trained, accurate plot summaries can be generated by inputting the plot text to be processed into the trained summary generation model.

[0199] In one embodiment, the plot text processing method of the present application is applied to the plot understanding of a film and television script, and the plot text processing method of the present application is described.

[0200] The inventor believes that the plot understanding for a film and television script is an important part of intelligent script understanding. Since each script of a film and television drama has several thousand to tens of thousands of scenes, and each scene has dozens to thousands of words, manual reading is inefficient and brings difficulties to script auditing. Automatic and rapid understanding of the plot of each scene based on artificial intelligence helps film and television auditing personnel quickly understand the plot and evaluate the script to evaluate whether a script has the value of shooting or related parties understand the plot and grasp the scene trend to better perform each scene. Using artificial intelligence for automatic script plot understanding, the traditional method is to use a text summary extraction model to automatically extract plot summaries from each scene text of a script. However, the traditional method has the following problems: first, plot summary generation often involves six elements such as time, place, character, cause (motivation), process, and result. Different elements have different effects on plot summary generation, and conventional plot summary learning does not consider the auxiliary role of such information on the generation result. Second, the summary generation model learning is sensitive to data, and biased training data will make the model generate biased language when generating a plot summary, thereby affecting the generalization ability of the summary generation model to language changes.

[0201] To solve the above problems in summary generation, the present application proposes a plot text processing method, which includes extracting low-level information related to the summary ability that the model needs to learn finally (i.e., plot category), assigning hierarchical tasks at the model structure level, and predicting low-level information before the model predicts the summary. By incorporating more low-level information, the high-level plot summary is outputted, which allows the model to generate different high-level summary descriptions while maintaining the core low-level information, thereby ensuring the sensitivity of the model to low-level information, promoting the learning of different low-level summary characteristic information according to different low-level plot categories, and improving the effect of each level of summary.

[0202] It can be understood that the highlights of the plot text processing method of the present application are: first, the low-level basic information (plot category) is combined into the summary generation model for learning, which supports high-level plot summary generation. Second, based on the training method of plot summary hierarchical data, the plot summary is learned from different hierarchical balances, which reduces the influence of biased data and helps to improve the effect of the summary under different hierarchical plot categories. Third, it can support the hierarchical learning of multiple low-level information combined with plot summaries.

[0203] It can be understood that the summary generation model generated in the plot text processing method of the present application can be applied to at least the following scenarios:

[0204] One is the script summary: the summary generation model can generate a concise summary of the script, helping readers quickly understand the main content and plot of the script. For example, for the input script text, first divide it into multiple scene texts, then pass each scene through the trained summary generation model to generate the scene understanding result. Input the script and return the plot of each scene for relevant personnel to review.

[0205] 2 Content planning: during the planning stage of a movie, TV series or other media project, the summary generation model can help planners quickly understand the main content of a large number of scripts, thereby making effective screening and decision-making.

[0206] 3 Script editing: for script editors or screenwriters, the summary generation model can help them quickly understand the main content of the script, thereby making effective modifications and improvements.

[0207] In one embodiment, the cascade level information learning process of the plot text processing method of the present application in the training model is first described, as shown in Figure 10 The output plot summary process after event layering in the plot text processing method of the present application is shown in Figure 10 The low-level information (HL) output by the initial summary generation model (as shown in Figure 10 The predicted summary output by the initial summary generation model).

[0208] Further, as shown in Figure 11 The two levels of information are output by the model in sequence, i.e. low-level information (as shown in Figure 11 The category feature of the predicted category output by the initial summary generation model) and high-level information (as shown in Figure 11 The predicted summary output by the initial summary generation model), and each level is learned separately. The low-level information is learned using the HL task (low-level task) loss (calculate category prediction loss value), and the high-level information is learned using the HH task (high-level task) loss (calculate summary prediction loss value). As shown in Figure 11 The plot summary and plot category in the target training sample are input as supervision information during the training phase.

[0209] Further, considering that the high-level task can be associated with multiple low-level tasks, the plot text processing method in the present application involves multiple low-level tasks as cascade tasks, wherein for the high-level task, the low-level task can be classification of character relationships, plot types, etc. For example, as shown in Figure 12As shown, taking the example of associating two low-level tasks, the low-level information output by the model includes the category feature 1 of the predicted category and the category feature 2 of the predicted category, and the high-level information is the predicted summary. Each level is learned separately, and the low-level information is learned using the HL task (low-level task) loss (calculate the category prediction loss value), and the high-level information is learned using the HH task (high-level task) loss (calculate the summary prediction loss value). Among them, the plot summary, plot category 1 and plot category 2 in the target training sample are input as supervision information in the training stage.

[0210] It should be noted that the hierarchical information joint model learning process in the present application will generate low-level information from the word representation aggregation of the output predicted summary before the model outputs the predicted summary. The low-level information should include the low-level information required for the plot summary. Therefore, in the present application, low-level supervision (i.e. plot category in the target training sample) is used to improve the prediction effect of low-level category prediction, so that the high-level plot summary contains sufficient and consistent low-level information. When there are at least two low-level tasks, the prediction tasks of the plot category prediction and the plot summary prediction of the two levels need to be balanced.

[0211] For example, taking the plot summary extraction of the plot text and the classification method of classifying according to the relationship between the characters as an example, the relationship between the characters is the low-level information of the level, and the input of the plot text is generally in the form of dialogue. The task of plot summary extraction is to summarize the main meaning of the dialogue text. Considering that the focus of the summary output will be different under different low-level understanding, such as in different character relationships, the model's understanding of the text and the way of generating summaries will be different - the dialogue between lovers focuses on emotional relationships rather than flirting, and the dialogue between parents and children focuses on educational themes. The dialogue text produced by two different characters will produce summaries with different focuses, such as dialogue between lovers focusing on emotional relationships rather than flirting, and dialogue between parents and children focusing on educational themes. Therefore, the present application improves the summary effect by using low-level capabilities (character relationship prediction) to guide before outputting the summary. It can be understood that in addition to the relationship between the characters, other classification methods can also be used for classification, such as plot type, time background, plot occurrence background location, etc.

[0212] In one embodiment, the plot text processing method in the present application, when training the model, the initial summary generation model will generate low-level information (i.e. the category feature of the predicted category) by information extraction on the predicted vector before outputting high-level information (i.e. the predicted summary). The high-level and low-level tasks are learned simultaneously during the learning of the initial summary generation model. The following introduces the summary generation model obtained by training from the aspects of data preparation, model structure and training, and application.

[0213] In one embodiment, in terms of data preparation, a plurality of training samples need to be generated, the training samples including plot text, plot abstract, and at least one plot category to which the plot text belongs, each plot category being obtained by classifying the plot text according to a classification manner. The plot text is text data that can be directly obtained, and the plot abstract and the at least one plot category to which the plot text belongs need to be obtained by further processing the plot text.

[0214] The determination of the at least one plot category to which the plot text belongs will be described below by taking the classification manner of classifying according to character relationship as an example.

[0215] It can be understood that each episode of a film or television drama contains interactions (including dialogues) of specific characters, and different character relationships produce different dialogue tones. Therefore, it is necessary to extract the in-episode main character relationship information of each episode of plot text. The specific extraction manner will be described below by taking a pre-trained natural language model GPT (Generative Pre-Trained Transformer) model as an example. The specific extraction manner can be as follows:

[0216] 1. For each episode (i.e., plot text) of the script, the topk (e.g., 2) characters with the highest frequency of appearance are retained from the characters appearing in the episode (the “person” label in the episode lists the characters included in the episode).

[0217] 2. For the retained character list, a pair of characters is taken, and the GPT is asked about the character relationship to generate a model character relationship. The question sentence is: “According to the following script, what is the relationship between a and b, where the relationship includes boyfriend, girlfriend, son, daughter, father, mother, colleague, classmate, and cannot be inferred. Only answer the relationship you are sure of. When the relationship is not clear, answer that it cannot be inferred. The script is: xxx”. In this way, 0-n character relationships are generated for each episode (0 relationships when there are no topk characters or no characters in the episode, and multiple relationships when there are multiple topk characters).

[0218] For example, for the first three characters ABC, the characters are taken in turn as (B, A) (C, A), and (C, B), i.e., asking who the less important character is. If A is the main character, it is equivalent to asking who B is to the main character A, so as to find the relationship of the key character.

[0219] 3. Collect the above results as character relationship data.

[0220] 4、According to the above GPT result, multiple character relationships are generated, and there may be multiple relationships between two people. According to the multiple relationships between two people, the final relationship is generated. The highest frequency relationship is selected from the multiple relationships as the final relationship, such as two people may be "pursuer", "admirer", "lover", and "lover" at the same time, and the final relationship is confirmed as lover.

[0221] 5、Select the two main characters in the scene - the top two people with the highest frequency of appearance, and keep the relationship between the two as the final character relationship of the scene. (This step can be omitted, and at this time, topk=2 is required).

[0222] In one embodiment, the plot summary can also be extracted by means of a pre-trained natural language model. Specifically, a pre-trained natural language model (such as GPT4) can be used to ask for the scene plot, and the question statement is "The following is a script of a certain scene of a movie, please extract the scene plot summary, and do not make guesses, do not make abstract descriptions, the script is xxx". It should be noted that since the plot obtained at this time may have detail errors or character misplacement, manual correction is required, and after correction, the data of scene script (plot text) - scene summary (plot summary) is obtained.

[0223] In one embodiment, the classification is classified according to the relationship between the characters, and the plot category is a specific character relationship. The form of the training data can be scene script (plot text) - scene plot (plot text) - character relationship (plot category), and the specific data format can be:

[0224] {"id": "999", "conversations":[

[0225] {"from":"human","value":"The following is a script of a certain scene of a movie, please extract the scene plot summary, and do not make guesses, do not make abstract descriptions, the script is xxx"},

[0226] {"from":"gpt","value":"xxxx"}

[0227] ]}

[0228] Wherein, the script xxx is the specific movie script of a certain scene, that is, the plot text, and the returned xxxx is the scene plot in the above data preparation. Id is the serial number of the data. The text input from human to the input end of the model is expected to output the text from gpt, wherein the text from gpt is the supervised information of the training.

[0229] In one embodiment, regarding the model structure, such as Figure 13 As shown, the initial summary generation model includes a summary prediction network and a plot classification network (such as...). Figure 13 As shown, there are N plot classification networks corresponding to N classification methods (two plot classification networks are shown). The summary prediction network includes an encoding layer, multiple word prediction networks, and a fully connected layer. The word prediction network can specifically be a DecoderLayer (e.g., Figure 13 As shown, using a 32-word prediction network as an example (for illustration), as... Figure 13 As shown, the output of the fully connected layer is a prediction of token ids (unique identifiers corresponding to each word in the predefined dictionary) (e.g., Figure 13 The example shown is the token ID (i.e., word prediction). For a pre-defined dictionary with 64,000 words, the token IDs predicted by the fully connected layer are vectors of size 1x64,000, where each vector element represents the probability that the encoding of each word in the token IDs is 1. The DecoderLayer mainly includes a normalization layer, an attention layer, and a multilayer perceptron. For example, its main structure can be illustrated in Table 1.

[0230] Table 1

[0231]

[0232] It should be noted that, as Figure 13 As shown, to generate low-level category features of the predicted categories before the fully connected layers predict the predicted summary, a narrative classification network is introduced at the output of the first word prediction network. The narrative classification network includes pooling layers and fusion layers (specifically, fully connected layers). The fusion layer outputs the category features of the predicted categories. The pooling layer can use average pooling.

[0233] The following is combined Figure 13 The calculation of loss values ​​involved in several places in this application is explained.

[0234] like Figure 13As shown, the encoder 1 (encoding vector 1) is the supervision information for calculating the loss of HH, and the encoder 1 is obtained by encoding the plot summary in the target training sample as supervision information, that is, the encoder 1 is the encoding vector corresponding to the plot summary. It can be understood that the encoding involved in the present application is all encoding in units of words, that is, the mapping of text vocabulary to the preset dictionary token ids is generated, and the encoding vector corresponding to the plot summary is actually the word encoding corresponding to each word in the plot summary. The word encoding can be a one-hot encoding value, and the vector length is 1xNword, and Nword is the length of the vocabulary in the preset dictionary, such as 64000 in the above example. When calculating the loss, the token id (a numerical identifier used to represent a specific position (position information) in word encoding) predicted for each word position is used as the one-hot encoding value of the corresponding position of the supervision information to calculate the loss.

[0235] That is, for each word in the plot summary, the predicted loss value corresponding to the word needs to be calculated according to the word encoding corresponding to the word and the predicted probability of the word at the same position in the predicted summary, and then the predicted loss value of the summary is obtained according to the predicted loss value of each word in the plot summary. Figure 13

[0236] As shown in Figure 14 , the encoder 2 (encoding vector 2) is the supervision information for calculating the loss of HH, and the encoder 2 is obtained by encoding the plot category in the target training sample as supervision information, that is, the encoder 2 is actually the category feature of the plot category. The specific encoding method can be that the word encoding of each word in the plot category is first performed to obtain the word encoding vector corresponding to each word in the plot category, and then the word encoding vector corresponding to each word is pooled to obtain the category feature of the plot category. In the present embodiment, the pooling calculation can be maximum pooling calculation.

[0237] It should be noted that since the plot text processing method in the present application involves two tasks, including a low-level basic information task (in the present embodiment, specifically a classification task) and a high-level target task (in the present embodiment, specifically a summary generation character), the balance of the two tasks needs to be performed during training.

[0238] In one embodiment, on the basis of obtaining a plurality of training samples, the server can perform full-parameter training using the plurality of training samples, and the training task is to learn the three parts of the word prediction network, the full connection layer and the plot classification network in the summary prediction network. The specific training process is as follows:

[0239] ​1. Parameter initialization: the network structure weight of the pre-trained natural language model is used to initialize the initial summary generation model, and the newly added modules (merge, fusion, etc. extracted by HL, i.e. plot classification network) are initialized using Gaussian distribution of (0, 1) parameters. Among them, the pre-trained natural language model can be selected according to the actual application scene.

[0240] 2. Set learning parameters: the word prediction network, fully connected layer and plot classification network in the summary prediction network are learning parameters.

[0241] 3. Learning process: for the full sample (i.e. multiple training samples, the total number of samples is L), obtain the number N of classification methods used to determine the plot category to which the multiple training samples belong, and obtain the number bs of plot categories configured for a batch of training training samples. According to the number N of classification methods and the number bs of plot categories, determine the number of plot category extraction (assuming bs is greater than or equal to N, then the number of plot category extraction is bs / N) from the plot categories of each classification method under the multiple training samples. For each classification method, extract the target plot category from the plot categories of the classification method under the multiple training samples according to the plot category extraction number (bs / N), and extract the target training sample from the training sample containing the target plot category, to obtain the training sample M used for a batch of training. Then, each completion of J (J = L / M) batches represents a round (epoch) iteration, and each round iteration processes the full sample once, until the average prediction loss value under a certain epoch no longer decreases (i.e. the average value of the prediction loss value corresponding to each batch training no longer decreases). During each batch learning:

[0242] (1) Model forward: input the plot text in each target training sample into the model for forward calculation, generate prediction for plot category and plot summary, then calculate the task loss of HH task (summary extraction) and HL task (plot category prediction task), and the total loss.

[0243] (2) Model backward: the total loss returns to the network to calculate the gradient of each parameter of the network.

[0244] (3) Model parameter update: update each parameter according to the gradient of each parameter of the network.

[0245] Among them, the total loss is: Loss1 = Loss_HH+ a*Loss_HL.

[0246] Among them, Loss_HH is the HH task loss, i.e. the summary prediction loss value, Loss_HL is the HL task loss, i.e. the category prediction loss value, and a is the weight, which can take values such as 0.1, 0.2, etc.

[0247] The following describes each task loss and related loss value calculation.

[0248] wherein the HH task loss is the summary prediction loss, and in this application, since the HH task is summary extraction, the summary extraction loss-cross entropy loss is used, i.e., the classification loss value of each word in the predicted plot summary, and it should be noted that the class probability here comes from each word in the preset dictionary, and each word can be considered as a class. The formula for calculating the summary prediction loss value can be:

[0249] ;

[0250] wherein represents the summary prediction loss value of a single target training sample, and N represents the number of target training samples processed in batches, which in this embodiment can be M, represents the word encoding corresponding to each word in the plot summary, and for each word in the plot summary, the word encoding corresponding to the word is determined according to the position of the word in the preset dictionary, and specifically can be one-hot (one-hot encoding) label, when it is a certain word in the preset dictionary, the label of the position corresponding to the word in the preset dictionary in the word encoding is 1, and the labels of the other positions are 0, represents the predicted probability corresponding to each word in the predicted summary, i.e., in the case of as the actual probability distribution of the plot summary, the predicted probability distribution of the predicted summary is compared to determine the summary prediction loss value.

[0251] wherein the HL task loss is the category prediction loss, and since HL is a newly added low-level task in hierarchical learning for supporting the final high-level task HH, the HL task prediction is output before the output of HH. It should be noted that for each classification method, the plot category classified according to the classification method and the predicted category classified and predicted according to the classification method are used as the loss calculation parameters of the classification method, and the category features of the plot category and the predicted category in the loss calculation parameters are used to calculate the prediction loss value corresponding to the classification method, and the prediction loss values corresponding to each classification method are superimposed to obtain the category prediction loss value. This scheme uses a similarity loss for HL to measure the similarity between the low-level prediction and the low-level supervision information, and the similarity between the low-level task prediction and the true low-level information is improved through the loss, so that the model output contains more such low-level information. That is, for each classification method, a similarity loss can be used to calculate the prediction loss value corresponding to the classification method.

[0252] ​Specifically, the manner of calculating the prediction loss value corresponding to the classification manner can be: encoding the plot category in the loss calculation parameter to obtain a plot category feature corresponding to the plot category in the loss calculation parameter, calculating a similarity loss value between the plot category feature and a category feature of the predicted category in the loss calculation parameter, and taking the similarity loss value as the prediction loss value corresponding to the classification manner.

[0253] The similarity loss can be specifically a Contrastive loss, i.e., the similarity between the plot category feature and the category feature of the predicted category in the loss calculation parameter is measured by the Contrastive loss. In the embodiment, the input of the Contrastive loss is a sample pair composed of the plot category feature and the category feature of the predicted category in the loss calculation parameter, and the label is whether the sample pair belongs to the same category, i.e., whether the predicted categories represented by the plot category feature and the category feature of the predicted category in the loss calculation parameter are the same.

[0254] In a specific application, the formula for calculating the similarity loss value can be:

[0255] ;

[0256] wherein, is the abbreviation of the function , which represents the vector after the input is mapped, i.e., the plot category feature in the embodiment, is the vector after the input is mapped, i.e., the category feature of the predicted category in the embodiment, and both represent the plot text, 1{·} is an indicator function that returns 1 when the input is true and returns 0 otherwise, i.e., returns 1 when the plot category in the loss calculation parameter and the category feature of the predicted category represent the same predicted category, and returns 0 otherwise, and m is a pre-set hyperparameter that can be configured according to the actual application scenario, indicating that the distance between different categories should exceed this value.

[0257] In an embodiment, the application of the plot text processing method of the present application is exemplified.

[0258] As shown in Figure 14 , the plot text processing method of the present application can be applied to rapid understanding of scripts in an abstract service system, providing model services, and generating a concise summary of the script to help readers quickly understand the main content and plot of the script. As Figure 15As shown, for the input script text, the summary service system can first divide the script text into multiple session texts, and then input each session text into the trained summary generation model in turn to generate the plot understanding result of the session text, i.e., the session summary. After obtaining the full-plot summary, the summary service system outputs the full-plot summary to the front end, so that the front end displays the full-plot summary.

[0259] In one embodiment, the plot text processing method of the present application has the following advantages. First, a hierarchical plot summary generation method is provided by means of low-level information joint learning, which can fuse multiple low-level information to provide positive help for plot summary generation. Second, the hierarchical cascade structure can support multiple low-level task joint, so that the high-level task receives more low-level information before output. Finally, the plot text processing method of the present application provides a hierarchical language model structure design, which can be extended to any high-level task and low-level task (specifically, a classification task), thereby improving the high-level language task.

[0260] It should be understood that although each step in the flowchart involved in each embodiment as described above is displayed in sequence according to the direction of the arrow, these steps are not necessarily executed in sequence according to the direction of the arrow. Unless otherwise specified herein, the execution of these steps is not strictly limited in sequence, and these steps can be executed in other sequences. Moreover, at least part of the steps in the flowchart involved in each embodiment as described above can include multiple steps or stages, which are not necessarily executed at the same time, but can be executed at different times, and the execution sequence of these steps or stages is not necessarily sequential, but can be executed in rotation or alternation with at least part of other steps or steps or stages in other steps.

[0261] Based on the same inventive concept, the present application also provides a plot text processing device for implementing the plot text processing method described above. The implementation scheme for solving the problem provided by the device is similar to the implementation scheme described in the above method, so the specific limitations in one or more plot text processing device embodiments provided below can refer to the limitations of the plot text processing method described above, which will not be repeated here.

[0262] In one embodiment, as shown in Figure 16 A plot text processing device is provided, comprising: a sample acquisition module 1502, a quantity acquisition module 1504, an extraction quantity determination module 1506, a training sample extraction module 1508, and a training module 1510, wherein:

[0263] The sample obtaining module 1502 is configured to obtain a plurality of training samples, wherein each training sample comprises a plot text, a plot abstract, and at least one plot category to which the plot text belongs, and each plot category is obtained by classifying plot texts according to a classification manner.

[0264] The number obtaining module 1504 is configured to obtain a number of classification manners used to determine plot categories to which the plurality of training samples belong, and obtain a number of plot categories configured for training samples used for a batch of training.

[0265] The extraction number determining module 1506 is configured to determine, according to the number of classification manners and the number of plot categories, a plot category extraction number of plot categories extracted from plot categories of each classification manner of the plurality of training samples.

[0266] The training sample extracting module 1508 is configured to, for each classification manner, extract a target plot category from plot categories of the classification manner of the plurality of training samples according to the plot category extraction number, and extract a target training sample from training samples containing the target plot category, to obtain training samples used for the batch of training.

[0267] The training module 1510 is configured to train a plot abstract generation model according to the training samples used for the batch of training. The trained plot abstract generation model is configured to output a plot abstract corresponding to a plot text to be processed based on the input plot text to be processed.

[0268] In an embodiment, the extraction number determining module is further configured to determine a ratio between the number of plot categories and the number of classification manners, and use the ratio as the plot category extraction number of plot categories extracted from plot categories of each classification manner of the plurality of training samples.

[0269] In an embodiment, the training sample extracting module is further configured to, for each classification manner, extract a target plot category from plot categories of the classification manner of the plurality of training samples according to the plot category extraction number, obtain a number of training samples configured for the target plot category, and extract a target training sample from training samples containing the target plot category according to the number of training samples, to obtain the training samples used for the batch of training.

[0270] In an embodiment, the training sample extraction module is further configured to cluster the plot summaries in each of the training samples containing the target plot category to obtain a summary category to which the plot summary in each of the training samples containing the target plot category belongs, determine a summary category extraction quantity of the summary categories extracted from the summary categories to which the plot summaries belong according to a number of the training samples and a number of the summary categories to which the plot summaries belong, extract a target summary category from the summary categories to which the plot summaries belong according to the summary category extraction quantity, and extract the target training samples from the training samples containing the target plot category and the plot summaries of which belong to the target summary category to obtain the training samples used for the batch training.

[0271] In an embodiment, the training sample extraction module is further configured to extract a representation of the plot summary in each of the training samples containing the target plot category to obtain a representation of the plot summary in the training sample, and cluster the representations of the plot summaries in each of the training samples containing the target plot category to obtain a summary category to which the plot summary in each of the training samples containing the target plot category belongs.

[0272] In an embodiment, the training module is further configured to input the plot text in each of the target training samples into the initial summary generation model to perform plot summary prediction and plot category prediction to obtain a predicted summary and a category feature of at least one predicted category, each of the predicted categories being obtained by performing classification prediction on the plot text according to a classification manner, perform prediction loss calculation according to the predicted summary, the category feature of the at least one predicted category, the plot summary in the target training sample, and at least one plot category to which the plot text belongs to obtain a prediction loss value corresponding to the target training sample, and adjust parameters of the initial summary generation model according to the prediction loss value corresponding to each of the target training samples to obtain the summary generation model.

[0273] In an embodiment, the training module is further configured to calculate a summary prediction loss value according to the predicted summary and the plot summary in the target training sample, calculate a category prediction loss value according to the category feature of the at least one predicted category and the at least one plot category to which the plot text belongs, and aggregate the summary prediction loss value and the category prediction loss value to obtain the prediction loss value corresponding to the target training sample.

[0274] In an embodiment, the training module is further configured to, for each classification manner, take the plot category classified according to the classification manner and the category feature of the predicted category predicted according to the classification manner as a loss calculation parameter of the classification manner, calculate a prediction loss value of the classification manner according to the category feature of the plot category and the category feature of the predicted category in the loss calculation parameter, and obtain a category prediction loss value according to the respective prediction loss values of each classification manner.

[0275] In an embodiment, the training module is further configured to encode the plot category in the loss calculation parameter to obtain a plot category feature corresponding to the plot category in the loss calculation parameter, calculate a similarity loss value between the plot category feature and the category feature of the predicted category in the loss calculation parameter, and take the similarity loss value as the prediction loss value of the classification manner.

[0276] In an embodiment, the training module is further configured to adjust the parameters of the initial summary generation model according to the prediction loss value corresponding to each target training sample in the training samples used in a batch training to obtain a model after parameter adjustment, determine the number of batch trainings in a round of training according to the total number of training samples and the number of training samples used in the batch training, and perform multi-batch training on the model after parameter adjustment according to the number of batch trainings to obtain the summary generation model.

[0277] In an embodiment, the plot text processing apparatus further comprises a plot summary generation module, and the plot summary generation module is configured to obtain the plot text to be processed, input the plot text to be processed into the trained summary generation model, and obtain a plot summary corresponding to the plot text to be processed.

[0278] Each module in the plot text processing apparatus described above can be realized wholly or partially by software, hardware, and a combination thereof. Each module described above can be embedded in or independent of a processor in a computer device in a hardware form, or stored in a memory in a computer device in a software form, so as to be called and executed by a processor to perform the operations corresponding to each module.

[0279] In an embodiment, a computer device is provided, which can be a server or a terminal. Taking the case that the computer device is a server, an internal structure diagram of the computer device can be shown. The computer device includes a processor, a memory, an input / output interface (I / O) and a communication interface. The processor, the memory and the input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. The processor of the computer device is configured to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for running the operating system and the computer program in the non-volatile storage medium. The database of the computer device is configured to store training samples and other data. The input / output interface of the computer device is configured to exchange information between the processor and external devices. The communication interface of the computer device is configured to communicate with external terminals through a network connection. The computer program is executed by the processor to implement a plot text processing method.

[0280] Those skilled in the art can understand that, ​ The structure shown in the figure is only a block diagram of part of the structure related to the scheme of the present application, and does not constitute a limitation on the computer device to which the scheme of the present application is applied. The specific computer device can include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.

[0281] In an embodiment, a computer device is provided, which includes a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the steps in the above method embodiments.

[0282] In an embodiment, a computer readable storage medium is provided, which stores a computer program. The computer program is executed by a processor to implement the steps in the above method embodiments.

[0283] In an embodiment, a computer program product is provided, which includes a computer program. The computer program is executed by a processor to implement the steps in the above method embodiments.

[0284] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer readable storage medium, and when the computer program is executed, the processes of the above-mentioned embodiments of the methods can be included. Any reference to memory, database or other medium used in the embodiments provided in the present application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical storage, high-density embedded non-volatile memory, resistive memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. As an illustration but not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc. The database involved in the embodiments provided in the present application can include at least one of a relational database and a non-relational database. The non-relational database can include a distributed database based on a block chain, etc., without being limited thereto. The processor involved in the embodiments provided in the present application can be a general-purpose processor, a central processing unit, a graphics processing unit, a digital signal processor, a programmable logic device, a data processing logic device based on quantum computing, etc., without being limited thereto.

[0285] Any combination of the technical features of the above embodiments can be made. In order to make the description simple, all possible combinations of the technical features in the above embodiments are not described, however, as long as the combination of the technical features does not exist, it should be considered as the scope of the present application.

[0286] The above embodiments only express several implementation manners of the present application, and the description is more specific and detailed, but it should not be understood as a limitation on the scope of the patent of the present application. It should be pointed out that for ordinary skilled in the art, without departing from the concept of the present application, a number of modifications and improvements can be made, which are within the scope of protection of the present application. Therefore, the protection scope of the present application should be subject to the appended claims.

Claims

1. A scenario text processing method characterized by comprising: The method comprises: obtaining a plurality of training samples, wherein each training sample comprises a plot text, a plot abstract, and at least one plot category to which the plot text belongs, and each plot category is obtained by classifying the plot text according to a classification manner; obtaining a number of classification manners used for determining the plot categories to which the plurality of training samples belong, and obtaining a number of plot categories configured for a batch of training samples; determining, according to the number of classification manners and the number of plot categories, a number of plot categories to be extracted from the plot categories of each classification manner of the plurality of training samples; for each classification manner, extracting a target plot category from the plot categories of the classification manner according to the number of plot categories to be extracted, and extracting a target training sample from the training sample comprising the target plot category, to obtain training samples used for the batch of training; for each target training sample in the training samples used for the batch of training, inputting a plot text in the target training sample into an initial abstract generation model to perform plot abstract prediction and plot category prediction, to obtain a predicted abstract and category features of at least one predicted category, wherein each predicted category is obtained by classifying the plot text according to a classification manner; calculating an abstract prediction loss value according to the predicted abstract and a plot abstract in the target training sample; for each classification manner, taking the plot categories obtained by classifying the classification manner and the category features of the predicted categories obtained by classifying and predicting the classification manner as loss calculation parameters of the classification manner; calculating a prediction loss value corresponding to the classification manner according to the plot categories and the category features of the predicted categories in the loss calculation parameters, superimposing the prediction loss values corresponding to each classification manner to obtain a category prediction loss value; aggregating the abstract prediction loss value and the category prediction loss value to obtain a prediction loss value corresponding to the target training sample; adjusting parameters of the initial abstract generation model according to the prediction loss value corresponding to each target training sample in the training samples used for the batch of training, to obtain an abstract generation model; and the trained abstract generation model is used to output a corresponding plot abstract of a to-be-processed plot text based on the input to-be-processed plot text.

2. The method of claim 1, wherein, The determination, according to the number of classification manners and the number of plot categories, of the number of plot categories to be extracted from the plot categories of each classification manner of the plurality of training samples comprises: determining a ratio between the number of plot categories and the number of classification manners; taking the ratio as the number of plot categories to be extracted from the plot categories of each classification manner of the plurality of training samples.

3. The method of claim 1, wherein, The target plot category is extracted from the plot categories of the plurality of training samples under the classification manner according to the plot category extraction quantity, and the target training sample is extracted from the training sample containing the target plot category, to obtain the training sample used for a batch training, including: For each classification manner, the target plot category is extracted from the plot categories of the plurality of training samples under the classification manner according to the plot category extraction quantity, and the training sample quantity configured for the target plot category is obtained; The target training sample is extracted from the training sample containing the target plot category according to the training sample quantity, to obtain the training sample used for a batch training.

4. The method of claim 3, wherein, The target training sample is extracted from the training sample containing the target plot category according to the training sample quantity, to obtain the training sample used for a batch training, including: The plot abstracts in each training sample containing the target plot category are clustered to obtain the abstract category to which the plot abstract in each training sample containing the target plot category belongs; According to the training sample quantity and the number of abstract categories to which the plot abstracts belong, the abstract category extraction quantity of extracting an abstract category from the abstract categories to which the plot abstracts belong is determined; The target abstract category is extracted from the abstract categories to which the plot abstracts belong according to the abstract category extraction quantity, and the target training sample is extracted from the training sample containing the target plot category and the plot abstract belonging to the target abstract category, to obtain the training sample used for a batch training.

5. The method of claim 4, wherein, The plot abstracts in each training sample containing the target plot category are clustered to obtain the abstract category to which the plot abstract in each training sample containing the target plot category belongs, including: For each training sample containing the target plot category, the representation of the plot abstract in the training sample is extracted to obtain the representation of the plot abstract in the training sample; The representations of the plot abstracts of each training sample containing the target plot category are clustered to obtain the abstract category to which the plot abstract in each training sample containing the target plot category belongs.

6. The method of claim 1, wherein, The prediction loss value corresponding to the classification manner is calculated according to the class features of the plot category and the prediction category in the loss calculation parameter, including: The plot category in the loss calculation parameter is encoded to obtain the plot category feature corresponding to the plot category in the loss calculation parameter; The similarity loss value between the plot category feature and the class feature of the prediction category in the loss calculation parameter is calculated, and the similarity loss value is taken as the prediction loss value corresponding to the classification manner.

7. The method of claim 1, wherein, The method further comprises: obtaining a to-be-processed plot text; inputting the to-be-processed plot text into the trained summary generation model to obtain a plot summary corresponding to the to-be-processed plot text. The device comprises:

8. The method according to any one of claims 1 to 7, characterized in that, a sample obtaining module configured to obtain a plurality of training samples, wherein each training sample comprises a plot text, a plot summary, and at least one plot category to which the plot text belongs, and each plot category is obtained by classifying the plot text according to a classification manner; a number obtaining module configured to obtain a number of classification manners used to determine plot categories to which the plurality of training samples belong, and obtain a plot category number configured for training samples used in a batch training; a number determining module configured to determine, according to the number of classification manners and the plot category number, a plot category number of extracted plot categories from each plot category under each classification manner of the plurality of training samples; 9. A scenario text processing apparatus characterized by comprising: a training sample extracting module configured to, for each classification manner, extract, according to the plot category number of extracted plot categories, a target plot category from plot categories of the plurality of training samples under the classification manner, and extract, from the training samples containing the target plot category, a target training sample to obtain training samples used in a batch training; ​ ​ ​ ​ The training module is configured to, for each target training sample in the training samples used for the batch training, input a plot text in the target training sample into an initial summary generation model to perform plot summary prediction and plot category prediction, to obtain a predicted summary and category features of at least one predicted category, each predicted category being obtained by classifying the plot text according to a classification manner; calculate a summary prediction loss value according to the predicted summary and a plot summary in the target training sample; for each classification manner, use a plot category obtained by classifying the plot text according to the classification manner and the category features of the predicted category obtained by classifying prediction according to the classification manner as loss calculation parameters of the classification manner; calculate a prediction loss value of the classification manner according to the plot category and the category features of the predicted category in the loss calculation parameters, superimpose the respective prediction loss values of each classification manner to obtain a category prediction loss value; aggregate the summary prediction loss value and the category prediction loss value to obtain a prediction loss value corresponding to the target training sample; and adjust parameters of the initial summary generation model according to the prediction loss value corresponding to each target training sample in the training samples used for the batch training, to obtain a summary generation model. The trained summary generation model is configured to output a plot summary corresponding to an input plot text based on the plot text.

10. The scenario text processing apparatus according to claim 9, characterized by The extraction quantity determination module is further configured to determine a ratio between the number of plot categories and the number of classification manners; and use the ratio as a plot category extraction quantity for extracting plot categories from plot categories of each classification manner of the plurality of training samples.

11. The scenario text processing apparatus according to claim 9, characterized by The training sample extraction module is further configured to, for each classification manner, extract a target plot category from plot categories of the classification manner of the plurality of training samples according to the plot category extraction quantity, and obtain a training sample quantity configured for the target plot category; extract a target training sample from the training samples containing the target plot category according to the training sample quantity, to obtain training samples used for the batch training.

12. The scenario text processing apparatus according to claim 11, characterized by The training sample extraction module is further configured to cluster plot summaries in each training sample of the training samples containing the target plot category, to obtain a summary category to which a plot summary in each training sample of the training samples containing the target plot category belongs; determine a summary category extraction quantity for extracting summary categories from the summary categories to which the plot summaries belong according to the training sample quantity and a number of the summary categories to which the plot summaries belong; extract a target summary category from the summary categories to which the plot summaries belong according to the summary category extraction quantity, and extract a target training sample from the training samples containing the target plot category and having plot summaries belonging to the target summary category, to obtain training samples used for the batch training.

13. The scenario text processing apparatus according to claim 12, characterized by The training sample extraction module is further configured to: for each training sample containing the target plot category, perform plot summary representation extraction on a plot summary in the training sample, to obtain a plot summary representation of the plot summary in the training sample; and perform clustering on the plot summary representation of each training sample containing the target plot category, to obtain a summary category to which a plot summary in each training sample containing the target plot category belongs.

14. The scenario text processing apparatus according to claim 9, characterized by The training module is further configured to encode plot categories in the loss calculation parameters, to obtain plot category features corresponding to the plot categories in the loss calculation parameters. The training module is further configured to: calculate a similarity loss value between the plot category features and category features of predicted categories in the loss calculation parameters; and use the similarity loss value as a predicted loss value corresponding to the classification mode.

15. The scenario text processing apparatus according to claim 9, characterized by The training module is further configured to: adjust parameters of the initial summary generation model according to predicted loss values corresponding to each target training sample in a batch of training samples used for training; determine a batch training frequency of one round of training according to a total number of the plurality of training samples and a number of training samples used in the batch training; and perform multiple batch trainings on the model with adjusted parameters according to the batch training frequency, to obtain a summary generation model. The apparatus further includes a plot summary generation module configured to: obtain a plot text to be processed; and input the plot text to be processed into the trained summary generation model, to obtain a plot summary corresponding to the plot text to be processed.

16. The scenario text processing apparatus according to claim 9, characterized by The processor, when executing the computer program, implements the steps of the method of any one of claims 1 to 8. 17.A computer device, comprising a memory and a processor, wherein the memory stores a computer program, and the computer device is configured to perform the method according to any one of claims 1-16. The computer program, when executed by the processor, implements the steps of the method of any one of claims 1 to 8.

18. A computer readable storage medium having stored thereon a computer program, characterized in that, The computer program, when executed by the processor, implements the steps of the method of any one of claims 1 to 8.

19. A computer program product comprising a computer program, characterized in that, ​

Citation Information

Patent Citations

  • Language processing model training method and device, electronic equipment and storage medium

    CN112559673A

  • Training method and device of plot extraction model, equipment and storage medium

    CN117875392A