Text generation method, device, electronic device and storage medium
By searching the target entries in the database and combining the chapter summary and previous chapter content to generate the initial text, the problem of character personality and plot consistency when generating novels is solved, and the quality and authenticity of the generated text are improved.
Patent Information
- Application Number
- CN202410634412.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-05-21
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2044-05-21
AI Technical Summary
When generating a complete novel, the current large language model is prone to problems such as inconsistent character characters, contradiction in plot development, unreasonable characters and plots, and poor content consistency and consistency.
By searching the target entry in the database, the initial text of the to-be-generated chapter is generated based on the content of the target entry, the chapter summary of the chapter to be generated and the content information of the previous chapter. Target entries can characterize character information and set information, helping the big model maintain character consistency and understanding of story structure.
It improves the quality of the generated initial text, ensures that the characters are consistent in character and reasonable in plot, enhances the authenticity and credibility of the text, and improves compliance with the original outline.
Smart Images

Figure CN118468819B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of artificial intelligence technology, and particularly to fields such as large models, large language models, generative models, etc. More specifically, the present disclosure provides a text generation method, apparatus, electronic device, storage medium, and computer program product. Background Art
[0002] In recent years, with the continuous evolution and wide application of artificial intelligence (AI) technology, AI creation has become a hot topic attracting wide attention. The AI technology can be applied to the field of literary creation. The AI can quickly transform the impressions in the author's mind into a story with detailed plot, and can add new details to enrich the story content, thereby greatly improving the author's writing efficiency.
[0003] However, with the current technical level of large language models, it is still relatively difficult to use AI to generate a complete novel. During the generation process, situations where the character personalities, plot developments, and detailed descriptions in the context are contradictory are likely to occur, and at the same time, situations where the characters and plots are unreasonable are also likely to occur. In addition, the coherence and consistency between the generated content of each part are poor. Summary of the Invention
[0004] The present disclosure provides a text generation method, apparatus, electronic device, storage medium, and computer program product.
[0005] According to one aspect of the present disclosure, there is provided a text generation method, including: determining at least one target entry from a database according to the chapter summary of the to-be-generated chapter; generating an initial text of the to-be-generated chapter according to the entry content of each of the at least one target entry, the chapter summary of the to-be-generated chapter, and the content information of the previous chapter; and determining the main text of the to-be-generated chapter according to the initial text of the to-be-generated chapter.
[0006] According to another aspect of the present disclosure, there is provided a text generation apparatus, including: an entry determination module, an initial text generation module, and a main text determination module. The entry determination module is configured to determine at least one target entry from a database according to the chapter summary of the to-be-generated chapter. The initial text generation module is configured to generate an initial text of the to-be-generated chapter according to the entry content of each of the at least one target entry, the chapter summary of the to-be-generated chapter, and the content information of the previous chapter. The main text determination module is configured to determine the main text of the to-be-generated chapter according to the initial text of the to-be-generated chapter.
[0007] According to another aspect of the present disclosure, there is provided an electronic device, including: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method provided by the present disclosure.
[0008] According to another aspect of the present disclosure, there is provided a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause a computer to execute the method provided by the present disclosure.
[0009] According to another aspect of the present disclosure, there is provided a computer program product, including a computer program which, when executed by a processor, implements the method provided by the present disclosure.
[0010] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. Description of the Drawings
[0011] The drawings are used to better understand the solution and do not constitute a limitation to the present disclosure. Among them:
[0012] Figure 1 is a schematic diagram of an application scenario of a text generation method and apparatus according to an embodiment of the present disclosure;
[0013] Figure 2 is a schematic flowchart of a text generation method according to an embodiment of the present disclosure;
[0014] Figure 3 is a schematic flowchart of a text generation method according to another embodiment of the present disclosure;
[0015] Figure 4 is a schematic principle diagram of a text generation method according to an embodiment of the present disclosure;
[0016] Figures 5A - 5F is a schematic diagram of a prompt information template according to an embodiment of the present disclosure;
[0017] Figure 6 is a schematic structural block diagram of a text generation apparatus according to an embodiment of the present disclosure; and
[0018] Figure 7 is a structural block diagram of an electronic device for implementing the text generation method according to an embodiment of the present disclosure. Detailed Embodiments
[0019] The following describes exemplary embodiments of the present disclosure with reference to the accompanying drawings. Various details of the embodiments of the present disclosure are included to facilitate understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, descriptions of well-known functions and structures are omitted in the following description for clarity and conciseness.
[0020] In the technical solution of the present disclosure, the processing of the collection, storage, use, processing, transmission, provision, and disclosure of the user's personal information complies with the provisions of relevant laws and regulations and does not violate public order and good customs.
[0021] In the technical solution of the present disclosure, the authorization or consent of the user is obtained before obtaining or collecting the user's personal information.
[0022] In some embodiments, during the process of creating a novel, an outline can be generated using a large language model based on the input inspiration, and then an abstract corresponding to each chapter can be generated based on the outline, and then the text of each chapter can be generated in sequence according to the abstract. However, when adopting this solution, the following problems are likely to occur.
[0023] First of all, when generating a long novel, since each chapter is generated separately and there is a lack of overall control of the overall plot and characters, situations such as inconsistent character personalities and contradictory plot developments that violate common sense may occur between different chapters. This brings great difficulties to the later manual integration and modification.
[0024] Secondly, when the AI generates some fictional characters and events that do not exist, obvious hallucination phenomena may occur. For example, generating some unreasonable character backgrounds or describing some historical events that did not occur. Such generation results lacking common sense constraints will seriously affect the authenticity and credibility of the novel.
[0025] Furthermore, when using this platform, users usually input a short novel synopsis as the basis for AI generation. However, in the actual generation process, the AI has poor compliance with the original synopsis. For example, the coherence and consistency of the generated content in each part are poor. After generating multiple parts of content, the later-generated content seriously deviates from the expected direction and cannot fully conform to and be faithful to the requirements of the original synopsis. A large amount of manual modification and sorting are required, increasing the later workload.
[0026] The present disclosure aims to provide a text generation method, which can retrieve target entries from a database, and then generate an initial text for the to-be-generated chapter according to the entry content of each target entry and the chapter summary of the to-be-generated chapter. The target entries can represent character information and setting information. Among them, the character information can keep the character personalities in different chapters consistent, and the setting information can constrain the content of the generated initial text to improve the authenticity and credibility of the text. It can be seen that when generating the initial text through a large model, the target entries can improve the global control ability of the large model and the understanding ability of the story structure, thereby optimizing the quality of the generated initial text. In addition, when generating the initial text of the to-be-generated chapter, the content information of the previous chapter can also be considered, which can improve the coherence and consistency between the initial text of the to-be-generated chapter and the content of the previous chapter, and further improve the compliance with the original outline.
[0027] The technical solutions provided by the present disclosure will be elaborated in detail below with reference to the accompanying drawings and specific embodiments.
[0028] Figure 1 It is a schematic diagram of the application scenario of the text generation method and device according to an embodiment of the present disclosure.
[0029] It should be noted that Figure 1 The figure shown is only an example of the system architecture to which the embodiments of the present disclosure can be applied, to help those skilled in the art understand the technical content of the present disclosure, but it does not mean that the embodiments of the present disclosure cannot be used in other devices, systems, environments or scenarios.
[0030] As Figure 1 shown, the system architecture 100 according to this embodiment may include terminal devices 101, 102, 103, a network 104, and a server 105. The network 104 is used to provide a medium for communication links between the terminal devices 101, 102, 103 and the server 105. The network 104 may include various connection types, such as wired and / or wireless communication links, etc.
[0031] Users can use the terminal devices 101, 102, 103 to interact with the server 105 through the network 104 to receive or send messages, etc. The terminal devices 101, 102, 103 may be various electronic devices with a display screen and supporting web browsing, including but not limited to smart phones, tablet computers, laptop portable computers, and desktop computers, etc.
[0032] Server 105 may be a server that provides various services. For example, it can be a background management server (only as an example) that supports websites browsed by users using terminal devices 101, 102, and 103. The background management server can analyze and process data such as user requests received, and feedback the processing results (such as the text of the chapter to be generated based on the inspiration or chapter summary input by the user, etc.) to the terminal devices.
[0033] It should be noted that the text generation method provided by the embodiments of the present disclosure can generally be executed by server 105. Correspondingly, the text generation device provided by the embodiments of the present disclosure can generally be set in server 105. The text generation method provided by the embodiments of the present disclosure can also be executed by a server or a server cluster different from server 105 and capable of communicating with terminal devices 101, 102, 103 and / or server 105. Correspondingly, the text generation device provided by the embodiments of the present disclosure can also be set in a server or a server cluster different from server 105 and capable of communicating with terminal devices 101, 102, 103 and / or server 105. The text generation method provided by the embodiments of the present disclosure can also be executed by terminal devices. Correspondingly, the text generation device provided by the embodiments of the present disclosure can be set in terminal devices.
[0034] It should be understood that Figure 1 the numbers of terminal devices, networks, and servers in
[0035] Figure 2 are merely illustrative. According to the implementation requirements, there can be any number of terminal devices, networks, and servers.
[0036] As Figure 2 shown, the text generation method 200 can be used in scenarios such as generating novels and stories. The text generation method 200 can include operation S210 to operation S230.
[0037] In operation S210, at least one target entry is determined from the database according to the chapter summary of the chapter to be generated.
[0038] Exemplarily, taking the generation of a novel as an example, a novel can include multiple chapters, and each chapter can correspond to a chapter summary for introducing the main plot of this chapter. The chapter summary can be written according to actual needs or generated based on the processing of user input inspiration by a large model.
[0039] Exemplarily, a database can be pre-constructed, and the present embodiment does not limit the construction process of the database. For example, the database can be pre-configured according to actual needs, or the data in the database can be generated based on a large model. The data in the database is called an entry.
[0040] Exemplarily, the database includes multiple entries, and each entry may include an entry name and entry content.
[0041] For example, an entry may represent character information of a to-be-generated chapter. The character information may represent characters in a novel, such as the protagonist and supporting characters. When an entry represents character information, the entry name may include the main name (name), alias, etc. of the character, and the entry content may include a character introduction, which may include the character's personality traits and the relationships between characters, and the relationships may include relatives, lovers, friends, etc.
[0042] Again, for example, an entry may represent setting information of a to-be-generated chapter. The setting information may include, for example, the world view setting of the story. The world view setting may include, for example, some custom or general nouns and explanations of these nouns. For example, the setting information of a wuxia novel may include "sect", "weapon", "skill", etc. When an entry represents setting information, the entry name may include the name, alias, etc. of the setting noun, and the entry content may include an introduction of the setting noun.
[0043] Exemplarily, in the process of determining the target entry, for example, an entry that appears in the chapter summary of the to-be-generated chapter may be determined as the target entry. Again, for example, a pre-trained model may be used to calculate the similarity between the entry and the chapter summary of the to-be-generated chapter, and an entry with a higher similarity may be determined as the target entry.
[0044] In operation S220, an initial text of the to-be-generated chapter is generated according to the respective entry contents of at least one target entry, the chapter summary of the to-be-generated chapter, and the content information of the previous chapter.
[0045] For example, the content information of the previous chapter may include the plot, ending, etc. of the previous chapter.
[0046] For example, an initial text may be generated based on a large model, and the large model may be, for example, a large language model, a generative model, etc. For example, a prompt information template for generating the initial text, the respective entry contents of the target entry, the chapter summary of the to-be-generated chapter, and the content information of the previous chapter may be concatenated into input information, and the input information is input into the large model, and the large model outputs the initial text of the to-be-generated chapter.
[0047] In operation S230, the main text of the to-be-generated chapter is determined according to the initial text of the to-be-generated chapter.
[0048] For example, the initial text of the to-be-generated chapter may be used as the main text of the to-be-generated chapter. Again, for example, errors in the initial text of the to-be-generated chapter may be corrected, and the corrected text may be used as the main text of the to-be-generated chapter.
[0049] The text generation method provided in this embodiment can retrieve target entries in a database, and then generate the initial text of the to-be-generated chapter according to the entry content of each target entry and the chapter summary of the to-be-generated chapter. The target entries can represent character information and setting information. Among them, the character information can keep the character personalities of different chapters consistent, and the setting information can constrain the content of the generated initial text to improve the authenticity and credibility of the text. It can be seen that when generating the initial text through a large model, the target entries can improve the global control ability of the large model and the understanding ability of the story structure, thereby optimizing the quality of the generated initial text. In addition, when generating the initial text of the to-be-generated chapter, the content information of the previous chapter can also be considered, which can improve the coherence and consistency between the initial text of the to-be-generated chapter and the content of the previous chapter, and further improve the compliance with the original outline.
[0050] Figure 3 It is a schematic flowchart of a text generation method according to another embodiment of the present disclosure.
[0051] As Figure 3 shown, the text generation method 300 may further include operations S301 to S305, which may be executed before the operation of determining at least one target entry from the database.
[0052] In operation S301, an outline is generated according to the object input text.
[0053] For example, the object may be a user, and the object input text may be the inspiration text input by the user. The outline may include the chapter summaries, character information, and setting information of multiple chapters.
[0054] In the process of generating the outline, it may be generated based on a large model, and this process may use the prompt information template Prompt_1 for generating the outline. For example, the prompt information template Prompt_1 may include the role and task of the large model, and the task may include: generating multiple chapter summaries, character information, and setting information based on the text input by the user. The prompt information template Prompt_1 may also include a small number of examples for prompting to improve the generation quality. It can be understood that in other embodiments, the outline written by the user may also be directly obtained.
[0055] In operation S302, target data is generated according to the outline.
[0056] In some embodiments, the target data includes at least one entry, and each entry in the at least one entry includes an entry name and an entry content.
[0057] For example, data in the database can be generated based on a large model, and in this process, the prompt information template Prompt_2 for generating entries can be used. The chapter summaries of the above-mentioned multiple chapters, the role information of the multiple chapters, the setting information of the multiple chapters, and the prompt information template Prompt_2 can be concatenated into input information, and the input information can be input into the large model. The large model outputs target data, and the target data can be JSON data.
[0058] For example, the prompt information template Prompt_2 can, for example, include the role and tasks of the large model, and the tasks can include: generating entries that meet the format requirements based on the chapter summary, role information, and setting information.
[0059] In some other embodiments, the target data can further include: a first identifier indicating the starting position of the entry and a second identifier indicating the ending position of the entry. For example, "{" can be used as the first identifier and "}" can be used as the second identifier. Each entry can have a first identifier and a second identifier, or a category of entries can have a first identifier and a second identifier. Among them, multiple role entries are a category of entries, and multiple setting nouns are another category of entries.
[0060] In operation S303, according to the first identifier and the second identifier, determine whether the target data meets the predetermined format condition. If it meets, perform operation S304; if it meets, perform operation S305.
[0061] For example, under normal circumstances, the first identifier and the second identifier appear in pairs. If the first identifier and the second identifier appear in pairs, it means that there is no abnormality in the format of the entry. At this time, the detected entry can be directly added to the database. If the first identifier and the second identifier do not appear in pairs, it means that the format of the entry is abnormal, and there is a possibility that the entry generated by the large model is inaccurate. Therefore, the target data can be regenerated.
[0062] In operation S304, update the database according to the entry name and entry content of each entry in the target data.
[0063] Operation S305, regenerate the target data.
[0064] For example, the prompt information template Prompt_2 can be updated by adding the new natural language text "The previous entry was incorrect, regenerate" to the prompt information template Prompt_2, and then repeat the above operation S302 based on the updated prompt information template Prompt_2 to regenerate the target data.
[0065] It can be seen that in this embodiment, it is verified whether the target data meets the predetermined format condition. If the predetermined format condition is not met, the target data is regenerated. After the target data meets the predetermined format condition, the database is updated. Therefore, the accuracy of the entries in the database can be improved by detecting the format, avoiding the chaos of the entries in the database, and further ensuring the accuracy of the initial text generated subsequently.
[0066] In some embodiments, the process of determining whether the target data meets the predetermined format condition may include: for a plurality of characters from the first first identifier to the last second identifier, traversing the plurality of characters based on the order of the plurality of characters, and recording the quantitative relationship between the number of first identifiers and the number of second identifiers during the traversal process. Then determine whether the quantitative relationship is consistent with the predetermined quantitative relationship. If so, it is determined that the target data meets the predetermined format condition; otherwise, it is determined that the target data does not meet the predetermined format condition. In this embodiment, the first identifier and the second identifier are respectively set before and after the entry, and whether the target data meets the predetermined format condition is determined based on the quantitative relationship between the first identifier and the second identifier. Therefore, the accuracy of the verification result can be improved.
[0067] For example, the predetermined quantitative relationship includes: during the traversal process, the number of first identifiers that have been traversed is greater than or equal to the number of second identifiers that have been traversed; and after the traversal is completed, the total number of first identifiers and the total number of second identifiers are equal. In this embodiment, whether the first identifier and the second identifier appear in pairs is used as the basis for verifying whether the format of the target data is accurate. This verification method is simple and easy to operate, and can ensure the accuracy of the verification result.
[0068] For example, in the actual verification process, taking "{" as the first identifier and "}" as the second identifier as an example, the positions between the first "{" and the last "}" in the output string can be matched. If the match fails, the verification ends and "incorrect" is returned. If the match is successful, the string can be traversed. If "{" appears, the counter is incremented by one; if "}" appears, the counter is decremented by one. If the counter is less than 0 during the process, the verification ends and "incorrect" is returned. If the counter is 0 after the traversal is completed, it means that the verification passes; otherwise, it means that the verification fails and the target data needs to be regenerated.
[0069] According to another embodiment of the present disclosure, the method for determining at least one target entry from a database according to the chapter summary of the chapter to be generated includes: determining, as candidate entries, the entries in the database whose entry names match the chapter summary of the chapter to be generated, to obtain at least one candidate entry. Then, according to the matching level between each candidate entry and the chapter summary of the chapter to be generated, and the entry length of each candidate entry, the at least one candidate entry is sorted to obtain sorting information. Then, based on the sorting information and the number of characters in the entry content of each of the at least one candidate entry, at least one target entry is determined from the at least one candidate entry.
[0070] For example, retrieval can be performed based on rules, so as to retrieve a part of the entries from the database as candidate entries. The retrieval rules may include: determining whether the name of the entry appears in the chapter summary, and if it appears, matching the entry with the chapter summary and taking the entry as the target entry. Or the retrieval rules may include: the similarity between the entry and the chapter summary is greater than the similarity threshold.
[0071] For example, after obtaining the candidate entries, if the number of candidate entries is too large, it will cause the number of input texts of the large model to exceed the maximum character limit. Therefore, the candidate entries can be sorted and screened.
[0072] For example, the candidate entries can be sorted in the order of hits, where the entry name can include the main name and the alias, and the matching priority of the main name is higher than that of the alias. Another example is that the candidate entries can be sorted according to the entries, where the priority of the short entries is higher than that of the long entries. In this way, the entries with the main name matched and the shorter entries can be ranked in the front position.
[0073] After obtaining the sorting information, the cumulative length of the entry content of the candidate entries can be calculated according to the sorting. When the cumulative length exceeds the length threshold after adding the entry content of a certain candidate entry, the other entries after this candidate entry can be deleted, and the remaining candidate entries are taken as the target entries.
[0074] In this embodiment, relevant candidate entries are retrieved first, then the candidate entries are sorted, and some of the candidate entries are screened as the target entries. In this way, more useful entries for the chapter to be generated can be recalled from the database, thereby improving the quality of the generated initial text.
[0075] According to another embodiment of the present disclosure, the process of generating the initial text of the to-be-generated chapter may include: generating the initial text of the to-be-generated chapter based on a large model. This process may use the prompt information template Prompt_3 for generating the initial text. The content of each of the above at least one target entry, the chapter summary of the to-be-generated chapter, and the content information of the previous chapter may be concatenated with the prompt information template Prompt_3 as input information, and the input information is input into the large model, and the large model outputs the initial text. For example, the prompt information template Prompt_3 may include, for example, the role and task of the large model, and the task may include: generating the initial text of a chapter based on the target entry and the chapter summary. This embodiment does not limit the prompt information template Prompt_3.
[0076] According to another embodiment of the present disclosure, the process of determining the body text of the to-be-generated chapter based on the initial text of the to-be-generated chapter may include: determining the error information in the initial text of the to-be-generated chapter according to the chapter summary of the to-be-generated chapter, the initial text of the to-be-generated chapter, and the role information of multiple chapters. Then, according to the chapter summary of the to-be-generated chapter and the error information, the initial text of the to-be-generated chapter is modified to obtain the body text of the to-be-generated chapter.
[0077] In the process of determining the error information, for example, a large model may be used to locate the position of the error information. This process may use the prompt information template Prompt_4 for locating the error information. The chapter summary of the to-be-generated chapter, the initial text of the to-be-generated chapter, and the role information of multiple chapters in the novel may be concatenated with the prompt information template Prompt_4 as input information, and the input information is input into the large model, and the large model outputs the error information.
[0078] For example, the prompt information template Prompt_4 may include, for example, the role and task of the large model, and the task may include: correcting the initial text to the body text based on the role information, the chapter summary of the to-be-generated chapter, and the initial text.
[0079] In the process of locating the error information, the large model may locate the position and type of the error. The position of the error may be represented in the form of "<start>.....<end>". The error information includes: multiple start texts representing the start position of the error segment in the initial text (i.e., the above "<start>"), and multiple end texts representing the end position of the error segment (i.e., the above "<end>"). The other text between the start text and the end text may be replaced by ".....", which reduces the output length.
[0080] After obtaining the error information, the initial text of the to-be-generated chapter may be modified according to the chapter summary of the to-be-generated chapter and the error information to obtain the body text of the to-be-generated chapter.
[0081] For example, for the position of the output of the large language model, it can be first matched in the initial text, and the method of regular expression can be used for matching. The regular expression is: <start>.+<end>, where ".+" represents any character more than one. If the match is successful, the incorrect segment will be retained, otherwise it will be discarded. In this way, the incorrect segment representing the error content can be intercepted from the initial text. In addition, the error information here can also include the incorrect segment and related explanations.
[0082] After obtaining the error information, the error information and the initial text can be input into the large model to complete the revision. For example, the modification operation can be performed based on the large model, and in this process, the prompt information template Prompt_5 for modifying the initial text can be used. The content of each of the above at least one target entry, the chapter summary of the chapter to be generated, and the content information of the previous chapter can be concatenated with the prompt information template Prompt_3 as the input information, and this input information is input into the large model, and the large model outputs the main text. For example, the prompt information template Prompt_5 can include the role and task of the large model, and the task can include: correcting the initial text based on the error information. In some embodiments, instead of presenting the initial text to the user, the revised main text can be directly presented.
[0083] In this embodiment, after obtaining the initial text of the chapter to be generated, the error information in the initial text is first determined, and then the initial text is modified, so as to eliminate problems such as incorrect character relationships, logical errors, and plot errors in the initial text, and improve the accuracy and logic of the main text of the chapter to be generated.
[0084] According to another embodiment of the present disclosure, after obtaining the main text of the chapter to be generated, the database can be updated based on the main text.
[0085] For example, the database can be updated based on the large model, and in this process, the prompt information template Prompt_6 for updating can be used. The main text of the chapter can be concatenated with the prompt information template Prompt_6 as the input information, and this input information is input into the large model, and the large model outputs relevant data. For example, the prompt information template Prompt_6 can include the role and task of the large model, and the task can include: generating entries based on the main text.
[0086] In addition, after obtaining the relevant data, the relevant data can be further processed in a manner similar to that for processing the target data, such as verifying whether the format of the relevant data is correct, and then adding the entries representing the roles or settings in the relevant data to the database under the correct circumstances.
[0087] Figure 4 It is a schematic diagram of the text generation method according to the embodiments of the present disclosure.
[0088] As Figure 4 shown, in this embodiment, the text generation method generates the required text based on a large model and a dynamic knowledge base, and the generated text can be a novel. The text generation method mainly includes the following processes: generating an outline, initializing a database, generating an initial text, correcting the initial text, and updating the database.
[0089] In the process of generating the outline, an outline can be generated according to the object input text, and the object input text can represent the user's inspiration 410. The outline can include the chapter summaries of multiple chapters, character information 430, and setting information 440. The chapter summaries of multiple chapters include, for example, the chapter summary 421 of the first chapter, the chapter summary 422 of the second chapter, and the chapter summary 423 of the third chapter.
[0090] In the process of initializing the database, target data can be generated according to the outline. The target data includes at least one entry, and each entry includes an entry name and entry content. The entry can be added to the database to update the database.
[0091] Next, the main text of each chapter can be generated through the process of generating the initial text and the process of correcting the initial text.
[0092] For example, the initial text 431 of the first chapter can be generated first according to the chapter summary 421 of the first chapter and the entries in the database, and then the main text 441 of the first chapter can be generated. For subsequent chapters, the content of the next chapter can be generated according to the content information of the previous chapter. For example, taking the second chapter as an example, at least one target entry can be determined from the database according to the chapter summary 422 of the second chapter. According to the entry content of each of the at least one target entry, the chapter summary 422 of the second chapter, and the content information of the first chapter (such as the plot and ending of the first chapter) obtained based on the main text 441 of the first chapter, the initial text 432 of the second chapter is generated. Then, the initial text 432 of the second chapter is corrected to obtain the main text 442 of the second chapter. Next, similar to the process of generating the main text 442 of the second chapter, the initial text 433 of the third chapter can be generated according to the content information of the second chapter obtained by using the main text 442 of the second chapter, the entries in the database, and the chapter summary 423 of the third chapter, and then the main text 443 of the third chapter is obtained by modification. In addition, each time the main text of a new chapter is generated, the entries in the database can be updated.
[0093] In this embodiment, a series of processes such as generating an outline, initializing a database, generating initial text, correcting the initial text, and updating the database can be performed based on the short inspiration 410 provided by the user, so as to better extract key elements such as characters and setting nouns in the novel, thereby improving the compliance and logic of the generated content, further enhancing the quality and efficiency of text generation, and improving the user experience. In addition, the dynamically updated knowledge base can be opened for users to edit, enhancing the user's fine control ability over the novel's plot.
[0094] Figures 5A - 5F It is a schematic diagram of the prompt information template according to an embodiment of the present disclosure.
[0095] In this embodiment, operations such as generating an outline, initializing a database, generating initial text for the chapter to be generated, determining error information in the initial text, modifying the initial text based on the error information, and updating the database can be performed based on a large model.
[0096] Such as Figure 5A shown, an outline can be generated based on a large model. This process can use the prompt information template Prompt_1. The prompt information template Prompt_1 can include the role 511 and task 512 of the large model. The task 512 can include: generating multiple chapter summaries, character information, and setting information based on the text input by the user. The prompt information template Prompt_1 can also include a small number of examples 513 for prompting to improve the generation quality.
[0097] Such as Figure 5B shown, the data in the database can be initialized based on a large model. This process can use the prompt information template Prompt_2 for generating entries. The prompt information template Prompt_2 can, for example, include the role 521 and task 522 of the large model. The task 522 can include: generating entries that meet the format requirements based on the chapter summary. The prompt information template Prompt_2 can also include examples 523.
[0098] Such as Figure 5C shown, the initial text of the chapter to be generated can be generated based on a large model. This process can use the prompt information template Prompt_3 for generating initial text. The entry content of at least one of the above target entries, the chapter summary of the chapter to be generated, and the content information of the previous chapter can be concatenated with the prompt information template Prompt_3 as input information, and the input information is input into the large model, and the large model outputs the initial text.
[0099] The prompt information template Prompt_3 can, for example, include the role 531 and task 532 of the large model. The task 532 can include: generating the initial text of a chapter based on the target entry and the chapter summary. In addition, the task 532 can guide the large model to think step by step according to a predefined process in a chain of thought manner.
[0100] As Figure 5D shown, the error information in the initial text can be determined based on the large model. This process can use the prompt information template Prompt_4 for locating error information. The chapter summary of the chapter to be generated, the initial text of the chapter to be generated, and the character information of multiple chapters in the novel can be concatenated with the prompt information template Prompt_4 as input information, and this input information is input into the large model, and the large model outputs the error information. The prompt information template Prompt_4 can, for example, include the role 541 and task 542 of the large model. The task 542 can include: correcting the initial text to the main text based on the character information, the chapter summary of the chapter to be generated, and the initial text. In addition, the prompt information template Prompt_4 can also include an example 543.
[0101] As Figure 5E shown, the operation of modifying the initial text can be performed based on the large model. This process can use the prompt information template Prompt_5 for modifying the initial text. The content of each of the above at least one target entry, the chapter summary of the chapter to be generated, and the content information of the previous chapter can be concatenated with the prompt information template Prompt_3 as input information, and this input information is input into the large model, and the large model outputs the main text. The prompt information template Prompt_5 can, for example, include the role 551 and task 552 of the large model. The task 552 can include: correcting the initial text based on the error information.
[0102] As Figure 5F shown, the database can be updated based on the large model. This process can use the prompt information template Prompt_6 for updating. The main text of the chapter can be concatenated with the prompt information template Prompt_6 as input information, and this input information is input into the large model, and the large model outputs the relevant data including the entries. The prompt information template Prompt_6 can, for example, include the role 561 and task 562 of the large model. The task can include: generating entries based on the main text.
[0103] It should be noted that any one of the above prompt information templates can be pre-configured, and the prompt information template can be described in a natural language manner. In addition, Figures 5A - 5F the prompt information templates Prompt_1 to Prompt_6 shown are only examples, and other description methods can be adopted.
[0104] Figure 6It is a schematic structural block diagram of a text generation device according to an embodiment of the present disclosure.
[0105] As Figure 6 shown, the text generation device 600 may include a term determination module 610, an initial text generation module 620, and a body determination module 630.
[0106] The term determination module 610 is configured to determine at least one target term from a database according to the chapter summary of the chapter to be generated.
[0107] The initial text generation module 620 is configured to generate an initial text of the chapter to be generated according to the term content of each of the at least one target term, the chapter summary of the chapter to be generated, and the content information of the previous chapter.
[0108] The body determination module 630 is configured to determine the body of the chapter to be generated according to the initial text of the chapter to be generated.
[0109] According to an embodiment of the present disclosure, the term determination module includes: a candidate term determination sub-module, a sorting sub-module, and a target term determination sub-module. The candidate term determination sub-module is configured to determine, as candidate terms, the terms in the database whose term names match the chapter summary of the chapter to be generated, and obtain at least one candidate term; the sorting sub-module is configured to sort the at least one candidate term according to the matching level of each candidate term with the chapter summary of the chapter to be generated and the term length of each candidate term, and obtain sorting information; the target term determination sub-module is configured to determine at least one target term from the at least one candidate term based on the sorting information and the number of characters of the term content of each of the at least one candidate term.
[0110] According to an embodiment of the present disclosure, the body determination module includes: an information determination sub-module and a modification sub-module. The information determination sub-module is configured to determine error information in the initial text of the chapter to be generated according to the chapter summary of the chapter to be generated, the initial text of the chapter to be generated, and the role information of multiple chapters; the modification sub-module is configured to modify the initial text of the chapter to be generated according to the chapter summary of the chapter to be generated and the error information, and obtain the body of the chapter to be generated.
[0111] According to an embodiment of the present disclosure, the error information includes: multiple start texts representing the start positions of error segments in the initial text, and multiple end texts representing the end positions of the error segments.
[0112] According to an embodiment of the present disclosure, the database is obtained through the following modules: a target data generation module, a judgment module, and an update module. The target data generation module is configured to generate target data according to the chapter summaries of multiple chapters, the role information of multiple chapters, and the setting information of multiple chapters; wherein, the target data includes at least one entry, a first identifier indicating the starting position of the entry, and a second identifier indicating the ending position of the entry, and each entry in the at least one entry includes an entry name and entry content; the judgment module is configured to determine whether the target data meets a predetermined format condition according to the first identifier and the second identifier; the update module is configured to update the database according to the entry name and entry content of each entry in the target data in response to detecting that the target data meets the predetermined format condition.
[0113] According to an embodiment of the present disclosure, determining whether the target data meets a predetermined format condition according to the first identifier and the second identifier includes: traversing multiple characters from the first first identifier to the last second identifier based on the order of the multiple characters, and recording the quantitative relationship between the number of first identifiers and the number of second identifiers during the traversal process; and determining that the target data meets the predetermined format condition in response to detecting that the quantitative relationship is consistent with a predetermined quantitative relationship.
[0114] According to an embodiment of the present disclosure, the predetermined quantitative relationship includes: during the traversal process, the number of first identifiers that have been traversed is greater than or equal to the number of second identifiers that have been traversed; and after the traversal is completed, the total number of first identifiers and the total number of second identifiers are equal.
[0115] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, including at least one processor; and a memory communicatively connected to the at least one processor; the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the above-mentioned text generation method.
[0116] According to an embodiment of the present disclosure, the present disclosure also provides a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause a computer to execute the above-mentioned text generation method.
[0117] According to an embodiment of the present disclosure, the present disclosure also provides a computer program product, including a computer program, and the computer program implements the above-mentioned text generation method when executed by a processor.
[0118] Figure 7FIG. 0 shows a schematic block diagram of an exemplary electronic device 700 that can be used to implement embodiments of the present disclosure. The electronic device is intended to represent various forms of digital computers, such as, for example, laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, for example, personal digital processors, cellular telephones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely exemplary and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0119] As Figure 7 shown, the device 700 includes a computing unit 701 that can perform various appropriate actions and processes in accordance with a computer program stored in a read-only memory (ROM) 702 or a computer program loaded from a storage unit 708 into a random access memory (RAM) 703. In the RAM 703, various programs and data required for the operation of the device 700 can also be stored. The computing unit 701, the ROM 702, and the RAM 703 are connected to each other via a bus 704. An input / output (I / O) interface 705 is also connected to the bus 704.
[0120] A plurality of components in the device 700 are connected to the I / O interface 705, including: an input unit 706, such as, for example, a keyboard, a mouse, etc.; an output unit 707, such as, for example, various types of displays, speakers, etc.; a storage unit 708, such as, for example, a magnetic disk, an optical disk, etc.; and a communication unit 709, such as, for example, a network card, a modem, a wireless communication transceiver, etc. The communication unit 709 allows the device 700 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0121] The computing unit 701 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 701 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 701 executes the various methods and processes described above, such as the text generation method. For example, in some embodiments, the text generation method can be implemented as a computer software program tangibly embodied in a machine-readable medium, such as the storage unit 708. In some embodiments, part or all of the computer program can be loaded and / or installed onto the device 700 via the ROM 702 and / or the communication unit 709. When the computer program is loaded into the RAM 703 and executed by the computing unit 701, one or more steps of the text generation method described above can be executed. Alternatively, in other embodiments, the computing unit 701 can be configured to execute the text generation method by any other suitable means (e.g., by means of firmware).
[0122] Various embodiments of the systems and techniques described above in this document can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-chip (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a dedicated or general-purpose programmable processor, receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting the data and instructions to the storage system, the at least one input device, and the at least one output device.
[0123] The program code for implementing the methods of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when the program codes are executed by the processor or controller, the functions / operations specified in the flowcharts and / or block diagrams are implemented. The program codes can be executed entirely on the machine, partially on the machine, executed partially on the machine as an independent software package and partially on a remote machine, or executed entirely on a remote machine or server.
[0124] In the context of this disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0125] To provide for interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can also be used to provide for interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic, speech, or tactile input).
[0126] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of a communication network include: a local area network (LAN), a wide area network (WAN), and the Internet.
[0127] A computer system can include clients and servers. Clients and servers are generally remote from each other and typically interact through a communication network. The relationship of client and server arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other.
[0128] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved, and no limitations are imposed herein.
[0129] The above specific embodiments do not constitute a limitation on the protection scope of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub - combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure shall be included within the protection scope of this disclosure.
Claims
1. A text generation method, comprising: According to the chapter summary for introducing the main plot of the chapter to be generated, at least one target entry is determined from a database, comprising: for at least one candidate entry retrieved from the database, according to the matching level of each candidate entry with the chapter summary of the chapter to be generated and the entry length of each candidate entry, the at least one candidate entry is sorted to obtain sorting information; and based on the sorting information and the number of characters of the entry content of each of the at least one candidate entry, the at least one target entry is determined from the at least one candidate entry; Generate an initial text of the chapter to be generated according to the entry content of each of the at least one target entry, the chapter summary of the chapter to be generated, and the content information of the previous chapter; and The main text of the chapter to be generated is determined according to the initial text of the chapter to be generated.
2. The method according to claim 1, wherein: The step of determining at least one target entry from a database according to the chapter summary of the chapter to be generated further comprises: The entries in the database whose entry names match the chapter summary of the to-be-generated chapter are determined as candidate entries, thereby obtaining the at least one candidate entry.
3. The method according to claim 1, wherein: Determining the body of the chapter to be generated according to the initial text of the chapter to be generated includes: Determining error information in the initial text of the chapter to be generated according to the chapter summary of the chapter to be generated, the initial text of the chapter to be generated, and character information of multiple chapters; and According to the chapter summary of the chapter to be generated and the error information, the initial text of the chapter to be generated is modified to obtain the main text of the chapter to be generated.
4. The method according to claim 3, wherein: The error information includes: a plurality of start texts representing the start position of the error segment in the initial text, and a plurality of end texts representing the end position of the error segment.
5. The method according to any one of claims 1 to 4, wherein: The database is obtained by: Generate target data according to the chapter summaries of each of the multiple chapters, the character information of the multiple chapters, and the setting information of the multiple chapters; wherein the target data includes at least one entry, a first identifier indicating the starting position of the entry, and a second identifier indicating the ending position of the entry, and each entry in the at least one entry includes an entry name and an entry content; determining whether the target data satisfies a predetermined format condition according to the first identifier and the second identifier; and In response to detecting that the target data satisfies the predetermined format condition, the database is updated according to the entry name and entry content of each entry in the target data.
6. The method according to claim 5, wherein: The determining, based on the first identifier and the second identifier, whether the target data satisfies a predetermined format condition comprises: For a plurality of characters from a first first identifier to a last second identifier, traverse the plurality of characters based on the order of the plurality of characters, and record a quantitative relationship between the number of the first identifiers and the number of the second identifiers during the traversal process; and In response to detecting that the quantitative relationship is consistent with a predetermined quantitative relationship, it is determined that the target data satisfies the predetermined format condition.
7. The method according to claim 6, wherein: The predetermined quantitative relationship includes: During the traversal process, the number of the first identifiers that have been traversed is greater than or equal to the number of the second identifiers that have been traversed; and After the traversal is completed, the total number of the first identifiers is equal to the total number of the second identifiers.
8. A text generation device, comprising: An entry determination module, used to determine at least one target entry from a database according to a chapter summary for introducing a main plot of a chapter to be generated; An initial text generation module, used to generate an initial text of the chapter to be generated according to the entry content of each of the at least one target entry, the chapter summary of the chapter to be generated and the content information of the previous chapter; as well as A text determination module, used for determining the text of the chapter to be generated according to the initial text of the chapter to be generated; Wherein, the entry determination module includes: a sorting submodule, for sorting at least one candidate term retrieved from the database according to a matching level between each candidate term and the chapter summary of the to-be-generated chapter and a term length of each candidate term, to obtain sorting information; and The target term determination submodule is used to determine the at least one target term from the at least one candidate term based on the sorting information and the number of characters in the term content of each of the at least one candidate term.
9. The device according to claim 8, wherein: The entry determination module also includes: The candidate entry determination submodule is used to determine the entry whose entry name in the database matches the chapter summary of the to-be-generated chapter as a candidate entry, and obtain the at least one candidate entry.
10. The device according to claim 8, wherein: The text determination module comprises: an information determination submodule, for determining error information in the initial text of the chapter to be generated according to the chapter summary of the chapter to be generated, the initial text of the chapter to be generated, and character information of multiple chapters; and The modification submodule is used to modify the initial text of the chapter to be generated according to the chapter summary of the chapter to be generated and the error information to obtain the main text of the chapter to be generated.
11. The device according to claim 10, wherein: The error information includes: a plurality of start texts representing the start position of the error segment in the initial text, and a plurality of end texts representing the end position of the error segment.
12. The device according to any one of claims 8 to 11, wherein: The database is obtained through the following modules: A target data generation module, configured to generate target data according to the chapter summaries of each of the multiple chapters, the character information of the multiple chapters, and the setting information of the multiple chapters; wherein the target data includes at least one entry, a first identifier indicating the starting position of the entry, and a second identifier indicating the ending position of the entry, and each entry in the at least one entry includes an entry name and an entry content; a judgment module, configured to determine whether the target data satisfies a predetermined format condition according to the first identifier and the second identifier; and An updating module is used for updating the database according to the entry name and entry content of each entry in the target data in response to detecting that the target data meets the predetermined format condition.
13. The device according to claim 12, wherein: The determining, based on the first identifier and the second identifier, whether the target data satisfies a predetermined format condition comprises: For a plurality of characters from a first first identifier to a last second identifier, traverse the plurality of characters based on the order of the plurality of characters, and record a quantitative relationship between the number of the first identifiers and the number of the second identifiers during the traversal process; and In response to detecting that the quantitative relationship is consistent with a predetermined quantitative relationship, it is determined that the target data satisfies the predetermined format condition.
14. The device according to claim 13, wherein: The predetermined quantitative relationship includes: During the traversal process, the number of the first identifiers that have been traversed is greater than or equal to the number of the second identifiers that have been traversed; and After the traversal is completed, the total number of the first identifiers is equal to the total number of the second identifiers.
15. An electronic device, comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 7.
16. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to cause the computer to execute the method according to any one of claims 1 to 7.
17. A computer program product comprising a computer program, which, when executed by a processor, implements the method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Intelligent writing method, device, equipment and medium
CN115630640A