Method and device for generating long text

By grouping and optimizing the outline, the logic and coherence problems of long text generation in AI writing are solved, and the generated report has a clear structure and detailed content.

CN120633598APending Publication Date: 2025-09-12BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510768586.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-10
Publication Date
2025-09-12

AI Technical Summary

Technical Problem

Existing AI writing has difficulty maintaining content coherence and logic when generating long texts, and may ignore key information, resulting in uneven report quality.

Method used

By grouping materials to generate an initial outline, providing improvement suggestions and optimizing the outline based on the materials, using a large language model to gradually generate chapter content, and introducing the "convolution" concept for iterative optimization.

Benefits of technology

It improves the logical coherence and content depth of long text generation, avoids the confusion of ideas and disconnected content in traditional methods, and generates reports with reasonable structure and complete content.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120633598A_ABST
    Figure CN120633598A_ABST
Patent Text Reader

Abstract

The invention provides a method and device for generating a long text, and relates to the field of artificial intelligence, in particular to the field of artificial intelligence writing. According to the specific implementation scheme, materials used for generating a long text are grouped, an initial outline is generated according to the grouping result, the initial outline comprises a plurality of chapter catalogues, and one chapter catalogue is associated with one group of materials; generating improvement suggestions for the initial outline based on the materials; optimizing the initial outline based on the improvement suggestion to generate a new outline; and generating chapter content corresponding to each chapter directory based on the new outline. According to the embodiment, the long text with coherence and logicality can be written.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of artificial intelligence, in particular to the field of artificial intelligence writing, and specifically to a method and device for generating long texts. Background Art

[0002] Currently, common AI writing methods primarily rely on direct generation to create long texts. When faced with complex and lengthy writing tasks, these tools often struggle to maintain coherence and logic, potentially overlooking key information or straying from the topic during the generation process, resulting in inconsistent report quality. Summary of the Invention

[0003] The present disclosure provides a method, apparatus, device, storage medium, and computer program product for generating a long text.

[0004] According to a first aspect of the present disclosure, a method for generating a long text is provided, comprising: grouping materials used to generate the long text, and generating an initial outline based on the grouping results, wherein the initial outline includes a plurality of chapter directories, and each chapter directory is associated with a group of materials; generating improvement suggestions for the initial outline based on the materials; optimizing the initial outline based on the improvement suggestions to generate a new outline; and generating chapter content corresponding to each chapter directory based on the new outline.

[0005] According to a second aspect of the present disclosure, a device for generating a long text is provided, comprising: a grouping unit configured to group materials used to generate the long text, and generate an initial outline based on the grouping result, wherein the initial outline includes multiple chapter directories, and each chapter directory is associated with a group of materials; a suggestion unit configured to generate improvement suggestions for the initial outline based on the materials; an optimization unit configured to optimize the initial outline based on the improvement suggestions and generate a new outline; and a generation unit configured to generate chapter content corresponding to each chapter directory based on the new outline.

[0006] According to a third aspect of the present disclosure, an electronic device is provided, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute any one of the methods described in the first aspect.

[0007] According to a fourth aspect of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to enable the computer to execute any one of the methods according to the first aspect.

[0008] According to a fifth aspect of the present disclosure, a computer program product is provided, comprising a computer program, wherein the computer program implements any one of the methods according to the first aspect when executed by a processor.

[0009] The method and apparatus for generating long texts provided by the embodiments of the present disclosure can prevent similar content from appearing in different sections of an outline by grouping materials. Rather than directly generating long reports, a "divide and conquer" approach is adopted to generate reports in segments and then merge them. It does not rely on a specific model and does not require fine-tuning, but is a general process optimization solution. It solves the problems of directly generating long reports, such as poor logical coherence, lack of focus, and insufficient content depth.

[0010] It should be understood that the contents described in this section are not intended to identify the key or important features of the embodiments of the present disclosure, nor are they intended to limit the scope of the present disclosure. Other features of the present disclosure will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] The accompanying drawings are provided to facilitate a better understanding of the present invention and do not constitute a limitation of the present disclosure.

[0012] Figure 1 is an exemplary system architecture diagram in which an embodiment of the present disclosure may be applied;

[0013] Figure 2 is a flowchart of an embodiment of a method for generating a long text according to the present disclosure;

[0014] Figure 3 is a schematic diagram of an application scenario of the method for generating long text according to the present disclosure;

[0015] Figure 4 is a flowchart of another embodiment of the method for generating a long text according to the present disclosure;

[0016] Figure 5 is a structural diagram of an embodiment of a device for generating a long text according to the present disclosure;

[0017] Figure 6 It is a schematic diagram of the structure of a computer system of an electronic device suitable for implementing the embodiments of the present disclosure. DETAILED DESCRIPTION

[0018] The following description of exemplary embodiments of the present disclosure is made in conjunction with the accompanying drawings, including various details of the embodiments of the present disclosure to facilitate understanding. These details should be considered as merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications may be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.

[0019] Figure 1 An exemplary system architecture 100 is shown to which an embodiment of the method or apparatus for generating a long text of the present disclosure can be applied.

[0020] like Figure 1 As shown, system architecture 100 may include terminal devices 101, 102, 103, a network 104, and a server 105. Network 104 is a medium for providing communication links between terminal devices 101, 102, 103 and server 105. Network 104 may include various connection types, such as wired or wireless communication links or fiber optic cables.

[0021] Users can use terminal devices 101, 102, and 103 to interact with server 105 via network 104 to receive or send messages, etc. Various communication client applications can be installed on terminal devices 101, 102, and 103, such as writing applications, 3D video players, web browser applications, shopping applications, search applications, instant messaging tools, email clients, social platform software, etc.

[0022] Terminal devices 101, 102, 103 can be hardware or software. When terminal devices 101, 102, 103 are hardware, they can be various electronic devices with display screens and support text editing, including but not limited to smart phones, tablet computers, e-book readers, MP3 players (Moving Picture Experts Group Audio Layer III, Moving Picture Experts Group Audio Layer 3), MP4 (Moving Picture Experts Group Audio Layer IV, Moving Picture Experts Group Audio Layer 4) players, laptop computers and desktop computers, etc. When terminal devices 101, 102, 103 are software, they can be installed in the electronic devices listed above. It can be implemented as multiple software or software modules (for example, to provide distributed services), or it can be implemented as a single software or software module. No specific limitation is made here.

[0023] Server 105 may be a server that provides writing services. Users can send the writing topic and materials to the server through their terminal devices. The server generates an initial outline based on the materials, then optimizes the initial outline based on the materials to generate a new outline. Finally, the long text content is generated based on the new outline.

[0024] It should be noted that a server can be either hardware or software. When a server is hardware, it can be implemented as a distributed server cluster consisting of multiple servers, or as a single server. When a server is software, it can be implemented as multiple software programs or software modules (for example, multiple software programs or software modules used to provide distributed services), or as a single software program or software module. This is not specifically limited here. The server can also be a server in a distributed system, or a server integrated with blockchain. The server can also be a cloud server, or an intelligent cloud computing server or intelligent cloud host with artificial intelligence technology.

[0025] It should be noted that the long text generating method provided in the embodiments of the present disclosure is generally executed by the server 105 , and accordingly, the long text generating device is generally provided in the server 105 .

[0026] It should be understood that Figure 1 The number of terminal devices, networks and servers in the embodiment is merely illustrative. Any number of terminal devices, networks and servers may be provided as required.

[0027] Continue to refer Figure 2 , shows a process 200 of an embodiment of a method for generating a long text according to the present disclosure. The method for generating a long text includes the following steps:

[0028] Step 201: Group the materials used to generate the long text, and generate an initial outline based on the grouping results.

[0029] In this embodiment, the execution body of the method for generating a long text (eg Figure 1 The server (shown) can obtain the materials used to generate the long text and group them. The grouping can be based on various reasons, such as topic similarity, summary similarity, title similarity, etc. Materials can also be grouped using a clustering algorithm, whereby materials with similar content are grouped together. The purpose of grouping here is to consolidate similar content into a group, enhance the logic of the generated initial outline, and prevent similar content from appearing in different sections of the outline.

[0030] For each material group, a table of contents is generated through the large language model to obtain an initial outline. Among them, the initial outline includes multiple table of contents, and a table of contents is associated with a group of materials. Using the existing functions of the general large language model, it is only necessary to design a dedicated prompt word (prompt), which can include information such as role, task description, generated content requirements, generated content restrictions, etc. For example, if you are an xx industry researcher (role), please generate a long text outline (task description) based on the material. The outline needs to include: abstract, background technology, technical principles, experimental results, references (generated content requirements), and the number of cited references shall not exceed 10 (generated content restrictions).

[0031] Step 202: Generate improvement suggestions for the initial outline based on the material;

[0032] In this embodiment, in order to optimize the structure and content of the outline, it is necessary to provide some optimization suggestions based on the previously collected materials. This solution is based on the established outline structure, processing the materials one by one and generating an outline summary. At the same time, some improvement suggestions for the outline will be made, such as pointing out that a section lacks data support or can be split into multiple subsections.

[0033] Based on the generated multiple outline summaries, they are integrated one by one based on the large language model and output as outline suggestions. The outline suggestions here not only list the chapter names, but also explain what should be written in each section and which literature should be cited.

[0034] A large language model can be used to generate improvement suggestions for the initial outline based on the material. This requires the design of a special prompt. For example, based on the title of the material, improvement suggestions for the initial outline can be generated, including: materials that need to be supplemented, literature that needs to be cited, etc.

[0035] Optionally, a different language model than that used to generate the initial outline in step 201 can be used to generate improvement suggestions for the initial outline. Even if the same model is used, a dialogue can be re-established to design prompts for criticizing or evaluating the initial outline, such as "check whether the content of the initial outline matches the material" or "evaluate the initial outline."

[0036] Step 203: Optimize the initial outline based on the improvement suggestions to generate a new outline;

[0037] In this embodiment, the initial outline is optimized based on the improvement suggestions using a large language model to generate a new outline. A corresponding prompt is designed, for example, "Optimize the initial outline based on the improvement suggestions and regenerate the outline."

[0038] Step 204: Generate chapter content corresponding to each chapter list based on the new outline.

[0039] In this embodiment, a large language model is used to generate chapter content using structurally segmented materials. Design corresponding prompts. For example, generate chapter content based on the materials associated with each chapter table of contents. Note that references should be marked with citations.

[0040] The methods provided in the above embodiments of the present disclosure optimize traditional AI long text generation solutions by introducing a concept similar to "convolution" and iteratively generating text from a single chapter, enabling them to better handle long text generation tasks. This change in process enables large language models to organize content more systematically during long text generation, avoiding the problems of confusion and disjointed content that are prone to occur in traditional large language models during long text generation.

[0041] In some optional implementations of this embodiment, grouping materials for generating a long text includes: respectively determining titles of a plurality of materials for generating the long text; and grouping the plurality of materials based on similarities between the titles.

[0042] A large language model is used to extract the title of each source material from multiple sources. The similarity between the titles (cosine similarity or edit distance, for example) is then calculated, and sources with a title similarity greater than a predetermined value are grouped together. Grouping by title allows sources from the same group to be used to write chapters with the same title, ensuring consistency and logical flow.

[0043] In some optional implementations of this embodiment, an initial outline is generated based on the grouping results, including: determining an outline template from a candidate template set based on the subject of the long text, wherein the candidate templates in the candidate template set correspond to a table of contents for a subject; and generating an initial outline based on the grouping results and the outline template. A variety of outline templates can be pre-set for selection, and each outline template corresponds to a subject, such as scientific research papers, financial analysis, physics, news, etc. Different outline templates have specific tables of contents. For example, scientific research papers need to include experimental data and references, while news does not need to include experimental data and references. Selecting the corresponding outline template according to different topics can improve the accuracy of the generated outline, and constraining the large language model through the outline template can avoid the large language model from arbitrarily generating impractical outlines. This method introduces an optional manually constructed outline template to guide the generated initial outline to be more professional. This step can improve the professionalism of the outline generation for industry research reports and other vertical fields.

[0044] In some optional implementations of this embodiment, the method further includes: in response to not receiving a topic for the long text, determining a topic based on the source material. The topic for the long text may be user-entered. If the user does not enter a topic, the source material may be analyzed using a large language model to determine the topic. If the user forgets to enter a topic or the entered topic is inaccurate, the topic can be determined based on the source material. This improves the accuracy of outline generation.

[0045] In some optional implementations of this embodiment, suggestions for improving the initial outline are generated based on the materials, including: for each directory, an outline summary of the directory is generated based on a group of materials associated with the directory; and suggestions for improving the initial outline are generated based on the outline summaries of each directory. An outline summary can be extracted from a group of materials through a large language model. For example, a summary of each material can be generated separately, and then an outline summary can be generated based on the common points of multiple summaries. Suggestions for improving the initial outline can be generated based on the outline summaries of each directory through a large language model. For example, it can be pointed out that a section lacks data support or can be divided into multiple subsections. By generating outline summaries by chapter and generating improvement suggestions for the initial outline, the improvement suggestions can be more targeted and more accurate. This can improve the logic and coherence of generating long texts.

[0046] In some optional implementations of this embodiment, improvement suggestions for the initial outline are generated based on the outline summaries of each chapter, including: determining the required supporting information based on the outline summaries of each chapter; detecting whether the source material is missing material associated with the supporting information; and, in response to detecting the lack of material, outputting improvement suggestions for supplementing the material. The required supporting information can be determined based on the outline summaries of each chapter using a large language model. For example, if the outline summaries are experimental data, and the experimental data is not found in the existing source material, then an improvement suggestion for supplementing the experimental data is output. The source material can be proactively supplemented to make the generated long text more complete.

[0047] In some optional implementations of this embodiment, improvement suggestions for the initial outline are generated based on the outline summaries of each directory, including: for each directory, detecting whether the number of materials associated with the directory exceeds a predetermined threshold; for a target directory whose number of associated materials exceeds the predetermined threshold, outputting suggestion information for splitting the target directory.

[0048] If the amount of material is too large, you can split it into multiple sub-chapter directories. This will make the content of each chapter more evenly distributed.

[0049] In some optional implementations of this embodiment, optimizing the initial outline based on the improvement suggestions to generate a new outline includes: obtaining additional material associated with the argument information by receiving supplementary documents and / or searching for documents on the internet; and optimizing the initial outline based on the additional material to generate a new outline. Supplementary material can be provided in two ways. The initial outline is optimized based on the additional material using the method described above to generate a new outline. Optimizing the outline by supplementing material can improve the logic and coherence of long text content.

[0050] In some optional implementations of this embodiment, the initial outline is optimized based on the improvement suggestions to generate a new outline, including: grouping the materials associated with the target chapter directory to obtain a number of groups; creating a corresponding number of sub-chapter directories for the target chapter directory based on the number of groups, where each sub-chapter directory is associated with a group of materials. Chapters with excessive materials are split into sub-chapter directories to make the outline more layered and the generated long text content more logical.

[0051] In some optional implementations of this embodiment, generating improvement suggestions for the initial outline based on the outline summaries of each chapter includes generating at least one of the following suggestion information based on the outline summaries of each chapter: chapter title, chapter content requirements, and suggested references. This suggestion information can be generated using a large language model, and the content requirements can be directly defined in the prompt. Generating this suggestion information in stages and then generating the long text based on the suggestion information, rather than directly generating the content of the long text, can improve the quality of the long text.

[0052] In some optional implementations of this embodiment, generating the chapter content corresponding to each directory based on the new outline includes: based on the material associated with each directory, generating the chapter content for each directory from the leaf directory of the new outline upward step by step. Starting from the leaf node (the most granular chapter), the content of each chapter is generated upward step by step, with each paragraph only citing the material that should be cited in the paragraph, to ensure the accuracy and logic of the content.

[0053] In some optional implementations of this embodiment, the method further includes: merging the table of contents of chapters with similar contents; performing at least one of the following operations on the merged chapter contents: polishing, adding illustrations, and adding references. After all chapters are written, a chapter-level merge summary will be performed, and finally the overall polishing, illustrations, and reference processing will be performed to eventually generate a long article with a reasonable structure, complete content, and Markdown format. Illustrations can be added by searching for pictures, or pictures can be rendered using Markdown syntax. The text content generated by each chapter is optimized to improve the coherence and logic of the long text, and rich graphic content is generated to increase the attractiveness of the long text.

[0054] In some optional implementations of this embodiment, the material is obtained in the following manner: determining keywords in the subject of a long text; searching and crawling web page content based on the keywords; filtering out irrelevant information from the web page content to obtain filtered web page content; filtering out target content whose similarity to the subject is higher than a predetermined threshold from the filtered web page content; parsing the target content to obtain a title, abstract, and body; and generating material in a predetermined format based on the title, abstract, and body.

[0055] Materials can be collected through search or file upload. For search, users only need to provide the report title (or topic). This solution incorporates a series of processes to collect materials, including using a large language model to break down the title into key points to generate search keywords, obtaining relevant web pages through search engines, crawling web content, filtering irrelevant information, and calculating topic similarity to select high-quality materials. This generates structured materials for easy storage and use.

[0056] In some optional implementations of this embodiment, the material is obtained in the following ways: parsing the uploaded document to parse out the title, abstract and text; generating material in a predetermined format based on the title, abstract and text. If the method is to upload a file, the user not only needs to provide the title of the report, but also needs to provide some documents related to the report (papers, news, patents, reports, etc.). Then, each document will be parsed first, and the title, abstract, text, etc. will be analyzed through a large language model, and the format will be unified and converted into structured content. Whether it is through searching or uploading files, the final output is a piece of structured content, which is the prepared material. Generate structured material for easy storage and use.

[0057] Continue to see Figure 3 , Figure 3 This is a schematic diagram of an application scenario of the method for generating long text according to this embodiment. This embodiment can be applied to scenarios where high-quality long reports need to be generated, such as literature reviews in the field of academic research, industry analysis reports, business planning plans, technical development documents, etc. Taking the field of academic research as an example, researchers can use the method of this application to quickly generate a literature review report with a clear structure and detailed content. The specific application process is as follows:

[0058] Step 301: Collecting Materials: After researchers determine the research topic, they input the topic into a large language model that implements the method of this application. The large language model automatically collects materials, filtering out materials related to the topic from a large amount of academic literature, papers, and other resources.

[0059] Step 302, generating an initial outline: the large language model processes the collected materials, selects a suitable outline template according to the topic, and generates an initial outline based on the materials and the outline template.

[0060] Step 303, optimizing the initial outline: generating improvement suggestions for the initial outline based on the material, and optimizing the initial outline based on the improvement suggestions. The initial outline is further improved through convolutional structure optimization to obtain a new outline.

[0061] Step 304, generate chapter content: Based on the optimized new outline, the large language model generates the content of the literature review paragraph by paragraph, elaborating on the research background, current status, main results and existing problems of the topic, and finally generates a complete literature review report.

[0062] Furthermore, this application can also be applied in the business sector, helping companies quickly develop business plans. For example, when developing a marketing plan, a company can use this method to collect market research data, competitor information, and other materials to generate a complete plan that includes market analysis, target market positioning, marketing strategy development, budget allocation, and other content, thereby improving work efficiency and the quality of the plan.

[0063] Further references Figure 4 , which shows a process 400 of another embodiment of a method for generating a long text. The process 400 of the method for generating a long text includes the following steps:

[0064] Step 401: Group the materials used to generate the long text, and generate an initial outline based on the grouping results;

[0065] Step 402: Generate improvement suggestions for the initial outline based on the material;

[0066] Step 403: Optimize the initial outline based on the improvement suggestions to generate a new outline;

[0067] Steps 401-403 are substantially the same as steps 201-203, and therefore are not described in detail.

[0068] Step 404: Generate suggestion information for a new outline based on the material;

[0069] In this embodiment, the process is basically the same as step 402, except that the generation of improvement suggestions for the initial outline is replaced by the generation of suggestion information for the new outline, that is, the suggestion information is iteratively optimized.

[0070] Step 405 : Optimize the new outline based on the suggestion information to generate an updated new outline.

[0071] In this embodiment, the optimization method is the same as step 403, that is, iterative optimization of the outline.

[0072] Steps 404-405 may be repeatedly executed to perform multiple rounds of iterations on the generation of the outline.

[0073] Step 406: Generate chapter content corresponding to each chapter list based on the updated new outline.

[0074] Step 406 is substantially the same as step 204 and thus will not be described in detail.

[0075] The method provided by the above-mentioned embodiment of the present disclosure performs multiple rounds of iterations on the generation of the outline, and eventually generates a new outline. Each round of optimization is like a convolution kernel sweeping through the structure diagram, and finally leaving behind the version with the highest information density and the clearest logic. After multiple rounds of optimization, some similar content may be merged and redundant parts may be deleted to make the outline more refined and reasonable. Of course, based on the final effect, the number of iterations here can be 1 or N times.

[0076] Further references Figure 5 As an implementation of the methods shown in the above figures, the present disclosure provides an embodiment of a device for generating a long text. Figure 2 Corresponding to the method embodiment shown, the device can be specifically applied to various electronic devices.

[0077] like Figure 5 As shown, the long text generating apparatus 500 of this embodiment includes: a grouping unit 501, a suggestion unit 502, an optimization unit 503, and a generation unit 504. The grouping unit 501 is configured to group the materials used to generate the long text and generate an initial outline based on the grouping results, wherein the initial outline includes multiple chapter directories, each chapter directory is associated with a group of materials; the suggestion unit 502 is configured to generate improvement suggestions for the initial outline based on the materials; the optimization unit 503 is configured to optimize the initial outline based on the improvement suggestions and generate a new outline; and the generation unit 504 is configured to generate chapter content corresponding to each chapter directory based on the new outline.

[0078] In this embodiment, the specific processing of the grouping unit 501, the suggestion unit 502, the optimization unit 503 and the generation unit 504 of the long text generation device 500 can be referred to. Figure 2 This corresponds to step 201, step 202, step 203 and step 204 in the embodiment.

[0079] In some optional implementations of this embodiment, the grouping unit 501 is further configured to: respectively determine titles of multiple materials used to generate the long text; and group the multiple materials based on similarities between the titles.

[0080] In some optional implementations of this embodiment, the grouping unit 501 is further configured to: determine an outline template from a candidate template set based on the subject of the long text, wherein the candidate templates in the candidate template set correspond to a table of contents for a subject; and generate an initial outline based on the grouping result and the outline template.

[0081] In some optional implementations of this embodiment, the grouping unit 501 is further configured to: in response to not receiving the subject of the long text, determine the subject based on the material.

[0082] In some optional implementations of this embodiment, the suggestion unit 502 is further configured to: generate an outline summary of each chapter directory based on a set of materials associated with the chapter directory; and generate improvement suggestions for the initial outline based on the outline summary of each chapter directory.

[0083] In some optional implementations of this embodiment, the suggestion unit 502 is further configured to: determine the required argument information based on the outline summary of each chapter directory; detect whether the material is missing material associated with the argument information; and output improvement suggestions for supplementing the material in response to detecting the lack of material.

[0084] In some optional implementations of this embodiment, the suggestion unit 502 is further configured to: for each chapter directory, detect whether the number of materials associated with the chapter directory exceeds a predetermined number threshold; for a target chapter directory whose number of associated materials exceeds the predetermined number threshold, output suggestion information for splitting the target chapter directory.

[0085] In some optional implementations of this embodiment, the optimization unit 503 is further configured to: obtain new materials associated with the argument information by receiving supplementary documents and / or searching documents on the Internet; optimize the initial outline based on the new materials to generate a new outline.

[0086] In some optional implementations of this embodiment, the optimization unit 503 is further configured to: group the materials associated with the target chapter directory to obtain a number of groups; create a corresponding number of sub-chapter directories of the target chapter directory based on the number of groups, wherein one sub-chapter directory is associated with a group of materials.

[0087] In some optional implementations of this embodiment, the suggestion unit 502 is further configured to generate at least one of the following suggestion information based on the outline summary of each chapter directory: chapter name, chapter content requirements, and recommended references.

[0088] In some optional implementations of this embodiment, the generating unit 504 is further configured to generate chapter contents of each directory step by step, starting from the leaf directory of the new outline and moving upwards, based on the materials associated with each directory.

[0089] In some optional implementations of this embodiment, the generation unit 504 is further configured to: merge the chapter directories with similar contents; and perform at least one of the following operations on the merged chapter contents: polishing, adding illustrations, and adding references.

[0090] In some optional implementations of this embodiment, the device 500 also includes an acquisition unit, which is configured to: determine keywords in the subject of the long text; search and crawl web page content based on the keywords; filter out irrelevant information from the web page content to obtain filtered web page content; filter out target content whose similarity with the subject is higher than a predetermined threshold from the filtered web page content; parse the target content to obtain a title, abstract and text; and generate material in a predetermined format based on the title, abstract and text.

[0091] In some optional implementations of this embodiment, the apparatus 500 further includes an acquisition unit configured to: parse the uploaded document to obtain the title, abstract, and body; and generate material in a predetermined format based on the title, abstract, and body.

[0092] In some optional implementations of this embodiment, the apparatus 500 further includes an iteration unit configured to: generate suggestion information for a new outline based on the material; and optimize the new outline based on the suggestion information to generate an updated new outline.

[0093] In the technical solutions disclosed herein, the collection, storage, use, processing, transmission, provision and disclosure of user personal information involved comply with the provisions of relevant laws and regulations and do not violate public order and good morals.

[0094] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.

[0095] An electronic device comprises: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method described in process 200.

[0096] A non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause the computer to execute the method described in process 200.

[0097] A computer program product includes a computer program, which implements the method described in process 200 when executed by a processor.

[0098] Figure 6 A schematic block diagram of an example electronic device 600 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are provided as examples only and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0099] like Figure 6 As shown, the device 600 includes a computing unit 601, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 602 or a computer program loaded from a storage unit 608 into a random access memory (RAM) 603. Various programs and data required for the operation of the device 600 can also be stored in the RAM 603. The computing unit 601, the ROM 602, and the RAM 603 are connected to each other via a bus 604. An input / output (I / O) interface 605 is also connected to the bus 604.

[0100] Various components in device 600 are connected to I / O interface 605, including an input unit 606, such as a keyboard, mouse, etc.; an output unit 607, such as various types of displays, speakers, etc.; a storage unit 608, such as a magnetic disk, optical disk, etc.; and a communication unit 609, such as a network card, modem, wireless communication transceiver, etc. The communication unit 609 allows device 600 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0101] The computing unit 601 can be a variety of general and / or special processing components with processing and computing capabilities. Some examples of the computing unit 601 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units that run machine learning model algorithms, digital signal processors (DSPs), and any appropriate processors, controllers, microcontrollers, etc. The computing unit 601 performs the various methods and processes described above, such as the long text generation method. For example, in some embodiments, the long text generation method can be implemented as a computer software program that is tangibly contained in a machine-readable medium, such as a storage unit 608. In some embodiments, part or all of the computer program can be loaded and / or installed on the device 600 via the ROM 602 and / or the communication unit 609. When the computer program is loaded into the RAM 603 and executed by the computing unit 601, one or more steps of the long text generation method described above can be performed. Alternatively, in other embodiments, the computing unit 601 can be configured to perform the long text generation method by any other appropriate means (e.g., by means of firmware).

[0102] Various embodiments of the systems and techniques described herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system-on-chip systems (SOCs), programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system that includes at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.

[0103] The program code for implementing the method of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device so that when the program code is executed by the processor or controller, the functions / operations specified in the flow chart and / or block diagram are implemented. The program code can be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0104] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in conjunction with an instruction execution system, device or equipment. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium can include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0105] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0106] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer having a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), and the Internet.

[0107] A computer system may include a client and a server. The client and server are generally remote from each other and typically interact through a communication network. The client-server relationship arises through computer programs running on the respective computers and having a client-server relationship with each other. The server may be a cloud server, a server in a distributed system, or a server integrated with a blockchain.

[0108] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in this disclosure can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved. This is not a limitation herein.

[0109] The above specific embodiments do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art will appreciate that various modifications, combinations, sub-combinations, and substitutions may be made based on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure shall be included within the scope of protection of this disclosure.

Claims

1. A method for generating a long text, comprising: Grouping materials used to generate a long text, and generating an initial outline based on the grouping results, wherein the initial outline includes a plurality of chapter directories, and each chapter directory is associated with a group of materials; generating improvement suggestions for the initial outline based on the material; Optimizing the initial outline based on the improvement suggestions to generate a new outline; Generate chapter content corresponding to each chapter directory based on the new outline.

2. The method according to claim 1, wherein The material grouping to be used to generate the long text includes: Determine the titles of multiple materials used to generate long texts respectively; The plurality of materials are grouped based on similarities between titles.

3. The method according to claim 1, wherein Generating an initial outline according to the grouping results includes: Determining an outline template from a candidate template set based on a theme of the long text, wherein the candidate templates in the candidate template set correspond to a table of contents of a theme; An initial outline is generated based on the grouping result and the outline template.

4. The method according to claim 3, wherein: The method further comprises: In response to not receiving the subject of the long text, the subject is determined based on the material.

5. The method according to claim 1, wherein Generating improvement suggestions for the initial outline based on the material includes: For each directory, generate an outline summary of the directory based on a group of materials associated with the directory; Proposals for improving the initial outline are generated based on the outline summaries of each chapter directory.

6. The method according to claim 5, wherein: The outline summary based on each chapter directory generates improvement suggestions for the initial outline, including: Determine the required supporting information based on the outline summary of each chapter; Detecting whether the material is missing material associated with the argument information; In response to detecting the lack of material, an improvement suggestion for supplementing the material is output.

7. The method according to claim 5, wherein: The outline summary based on each chapter directory generates improvement suggestions for the initial outline, including: For each chapter directory, checking whether the number of materials associated with the chapter directory exceeds a predetermined number threshold; For a target chapter directory whose number of associated materials exceeds a predetermined number threshold, output suggestion information for splitting the target chapter directory.

8. The method according to claim 6, wherein: The step of optimizing the initial outline based on the improvement suggestions to generate a new outline includes: Acquiring new materials related to the argument information by receiving supplementary documents and / or searching documents on the Internet; The initial outline is optimized based on the newly added material to generate a new outline.

9. The method according to claim 7, wherein: The step of optimizing the initial outline based on the improvement suggestions to generate a new outline includes: Grouping the materials associated with the target chapter directory to obtain the number of groups; A corresponding number of sub-chapter directories of the target chapter directory are created according to the number of groups, wherein one sub-chapter directory is associated with a group of materials.

10. The method according to claim 5, wherein The outline summary based on each chapter directory generates improvement suggestions for the initial outline, including: Generate at least one of the following suggested information based on the outline summary of each chapter: chapter name, chapter content requirements, and recommended references.

11. The method according to claim 1, wherein Generating chapter contents corresponding to each chapter directory based on the new outline includes: Based on the materials associated with each chapter directory, the chapter contents of each chapter directory are generated step by step starting from the leaf chapter directory of the new outline.

12. The method according to claim 11, wherein The method further comprises: Merge chapters with similar content; Perform at least one of the following operations on the merged chapter content: polish, add illustrations, and add citations.

13. The method according to claim 1, wherein The materials are obtained in the following ways: determining keywords in the subject of the long text; Search and crawl web page content based on the keywords; filtering out irrelevant information from the webpage content to obtain filtered webpage content; Filtering target content whose similarity to the subject is higher than a predetermined threshold from the filtered web page content; Parsing the target content to extract the title, abstract and body; A material in a predetermined format is generated based on the title, the abstract, and the main text.

14. The method according to claim 1, wherein The materials are obtained in the following ways: Parse the uploaded document to extract the title, abstract and body; A material in a predetermined format is generated based on the title, the abstract, and the main text.

15. The method according to any one of claims 1 to 14, wherein Before generating chapter contents corresponding to each chapter list based on the new outline, the method further includes: generating suggestion information for the new outline based on the material; The new outline is optimized based on the suggestion information to generate an updated new outline.

16. A device for generating a long text, comprising: a grouping unit configured to group materials used to generate a long text and generate an initial outline according to the grouping result, wherein the initial outline includes a plurality of chapter directories, and each chapter directory is associated with a group of materials; a suggestion unit configured to generate improvement suggestions for the initial outline based on the material; an optimization unit configured to optimize the initial outline based on the improvement suggestion to generate a new outline; The generating unit is configured to generate chapter contents corresponding to each chapter list based on the new outline.

17. An electronic device comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 15.

18. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to cause the computer to execute the method according to any one of claims 1-15.

19. A computer program product comprising a computer program, which, when executed by a processor, implements the method according to any one of claims 1 to 15.