Knowledge enhancement and context semantic coherence report writing generation method and system based on large model
By using real-time information collection, personalized style design, and contextual coherence optimization, the problems of information lag, insufficient personalization support, and text generation length limitations in large model report writing have been solved, enabling the generation of efficient, personalized, and logically rigorous long reports.
Patent Information
- Application Number
- CN202510811040.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-17
- Publication Date
- 2025-11-07
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing large-model-based report writing technologies have shortcomings in information updates, contextual coherence, and personalization support, resulting in low-quality generated content that fails to meet diverse user needs. Furthermore, the limited text length affects writing efficiency and widespread application.
By introducing a knowledge enhancement module to collect multi-source heterogeneous data in real time, a personalized writing module to design prompt words, and combining a contextual semantic coherence optimization module and a module to overcome the text generation length limit, dynamic information integration, personalized generation, and logically rigorous long text generation are achieved.
It achieves timeliness, personalization, and contextual coherence in the generated content, breaks through the text generation length limit, significantly improves the efficiency and quality of report writing, and adapts to the diverse needs of different users.
Smart Images

Figure CN120911587A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application belongs to the technical field of artificial intelligence, and in particular relates to a knowledge enhancement and context semantic coherence report writing generation method and system based on a large model. BACKGROUND
[0002] In today's information age, report writing is not only a basic skill, but also a task that needs to be completed efficiently and with high quality. With the rapid development of large model technology, people's access to information channels is increasingly rich. How to use these massive information resources to improve the quality and efficiency of report writing has become the focus of attention of many researchers and practitioners. Existing large model-based writing has been widely used in news reporting, academic research, business reports and other fields, and its technology is constantly maturing, but there are still some limitations, such as serious hallucination, lack of up-to-date background knowledge, poor context coherence, and text generation length restrictions.
[0003] In view of the above problems, some improved large model-based report writing solutions have emerged. These solutions aim to improve the quality, efficiency and relevance of text generation. Rule-based writing solutions rely on predefined templates and grammar rules to construct sentences and paragraphs. In this way, it can ensure that the generated text meets specific format requirements and is grammatically correct. Large model fine-tuning technology can adapt large models to specific domain or type writing tasks, further improving the professionalism and accuracy of generated content. Integrating external data sources to enhance timeliness, in order to make the generated content more timely and relevant, some advanced writing solutions begin to integrate real-time updated data sources such as news websites, social media platforms, etc. This not only allows you to quickly capture the latest events and developments, but also provides users with the freshest information materials to help them create articles that keep up with the times.
[0004] Although existing technologies have improved the effectiveness of large model-based report writing to some extent, there are still the following main problems: lack of up-to-date background knowledge, unable to obtain and integrate the latest information resources in real time; lack of personalized support, unable to generate report writing according to the specific needs and style preferences of users, difficult to meet the diverse writing needs of different users; lack of context coherence, difficult to ensure logical connection between different parts of the generated document; text generation length restrictions, most large models have certain restrictions on the length of text generated at a time, which makes it difficult to generate longer reports.
[0005] Although the writing systems in the prior art have improved writing efficiency to some extent, there are still some problems that not only affect the quality of the generated content, but also limit the widespread application and development potential of large model-based writing. Specifically, the main defects and problems of existing technology can be summarized as follows:
[0006] (1) Information update lag
[0007] Existing large model report writing systems have a significant lag in information updating, mainly relying on pre-built knowledge graphs or fixed data sets. These static resources are difficult to update in real time, resulting in generated content that may not reflect the latest information and trends. This lag is particularly evident in rapidly developing fields such as news reporting, technology trends, and policy changes, making articles unable to meet readers' demand for the latest information. In addition, although modern information retrieval technology can provide a large amount of data, it lacks an efficient filtering mechanism, resulting in search results containing a large amount of irrelevant or low-quality information, increasing the workload of manual filtering for users and reducing writing efficiency. The problem of information overload further exacerbates this challenge, as too much information makes it difficult for users to distinguish valuable content, affecting the overall writing quality and efficiency.
[0008] (2) Lack of personalized support
[0009] Different users have different writing needs and style preferences, but existing writing solutions often use uniform templates or preset rules, lacking the ability to adjust flexibly and meet diverse individual needs. Weak customization is a major shortcoming, as they may only be suitable for specific types of documents (such as news releases or academic papers), and do not provide adequate support for other types of writing tasks (such as creative writing or business proposals). In addition, existing methods perform poorly in maintaining writing style consistency, for example, users may expect the generated article to have a specific tone, vocabulary selection, or sentence structure, but existing methods are difficult to achieve this. The lack of personalized feedback mechanisms is also a major problem, as there is no function to iteratively improve based on user feedback, so even if users express specific modification suggestions, the system cannot effectively respond, further weakening user experience. Therefore, the lack of personalized support limits the widespread application and development potential of large model writing.
[0010] (3) Lack of context coherence
[0011] Existing large models have difficulty maintaining logical coherence and consistency of context when generating long articles. This can result in loose article structure, content jumping, and inconvenience for readers to understand. Especially in complex documents with multiple chapters or paragraphs, the connection between different parts is not close enough, and there may even be repetition or contradiction, which seriously affects the overall quality of the article. In addition, the lack of semantic understanding is also a big problem. Although existing technologies can parse complex semantic structures, there may still be semantic breaks in the actual generation process, i.e. lack of natural transition between some sentences or paragraphs, affecting the fluency of the article. When it comes to professional terms or concepts in a particular field, the model may not be able to correctly understand and express these contents due to lack of sufficient background knowledge, resulting in a decrease in the professionalism and authority of the article.
[0012] (4) Text generation length limit
[0013] Most large models are designed to focus on short text tasks, with limited support for long and complex documents. This not only limits the length of text generated at a time, but also causes inconvenience for long report writing that needs to integrate a large amount of information, thereby affecting work efficiency. In addition, when dealing with long documents, in addition to the limitation of generation length, it is also necessary to consider how to effectively integrate information from different sources to ensure the consistency and accuracy of the content, which puts higher requirements on existing models. The lack of segmented generation and information integration mechanism makes users have to manually adjust and supplement, greatly reducing the convenience and efficiency of writing.
[0014] In view of the above analysis, the technical problems existing in the prior art that need to be solved urgently are:
[0015] The existing report writing technology based on large models has obvious shortcomings in information updating, context coherence and personalized support. These problems not only affect the quality of the generated content, but also limit the wide application and development potential of large models to some extent. Therefore, a more advanced and flexible method is needed to solve these problems and further improve the overall performance and efficiency of the system. SUMMARY
[0016] In view of the problems existing in the prior art, the present application provides a knowledge enhancement and context semantic coherence report writing generation method and system based on a large model.
[0017] The present application is implemented as follows: a knowledge enhancement and context semantic coherence report writing generation method based on a large model, the knowledge enhancement and context semantic coherence report writing generation method based on a large model, the method specifically comprises:
[0018] S1: The knowledge enhancement module receives user input topics or titles, obtains relevant information from designated data sources according to the input topics, performs data collection, calculates the similarity between the topics and the information data in the database and reorders them, selects the top 50 most relevant data, stores the filtered data for use by the large model, and provides the large model with the filtered data;
[0019] S2: The personalized style writing topic and article outline module collects user writing needs and style preferences, the large model generates high-quality article titles according to user input topics or keywords, combines user prompt word design and requirements, and generates standardized article outlines according to the topic and background knowledge base;
[0020] S3: The context semantic coherence optimization module divides the entire outline into multiple independent first-level titles according to the fine granularity of the first-level titles, and designs and writes content for each chapter using the large model prompt word technology, combined with the background knowledge base and the structure framework of the generated content.
[0021] S4: The text generation length limitation overcoming module iteratively generates content section by section according to multiple outline chapters.
[0022] Another object of the present application is to provide a large model-based knowledge enhancement and context semantic coherence report writing generation system, which specifically comprises:
[0023] The knowledge enhancement module is used to provide the latest data preparation for report writing;
[0024] The personalized style writing topic and article outline module is used to generate article titles and standardized article outlines that meet the format requirements according to user input topics or keywords, and provide personalized prompt word guidance for the large model to generate content;
[0025] The context semantic coherence optimization module is used to improve the logical coherence of long and complex report writing, make the generated report structure rigorous and clear, and meet the requirements of professional report writing;
[0026] The text generation length limitation overcoming module is used to iteratively generate content section by section according to multiple outline chapters.
[0027] Further, the knowledge enhancement module specifically implements the following process:
[0028] (1) Information source selection: Select multiple information sources such as Twitter, Google, WeChat public number, etc. to provide free information search platforms to all users;
[0029] (2) Information retrieval: Vectorize the collected data and store it in the vector database Milvus, and retrieve relevant background content from the selected information sources according to the user-provided question;
[0030] (3)Similarity calculation: vectorize the user-provided question, and calculate the similarity of the retrieval results using cosine similarity;
[0031] (4)Reordering: reorder the retrieval results according to the similarity, and select the top 50 most relevant results;
[0032] (5)Data integration: integrate the top 50 most relevant results into open source data background knowledge for use by large models.
[0033] Further, the personalized style writing topic and article outline module is implemented as follows:
[0034] (1)Prompt word design: design large model prompt words to guide the large model to generate article titles and outlines;
[0035] (2)Generate article title: generate article title according to user input;
[0036] (3)Generate official document outline: generate official document outline in strict accordance with markdown format.
[0037] Further, the length limitation overcoming text generation module is implemented as follows:
[0038] (1)Segment generation: generate content segment by segment according to the requirements of the large model prompt words;
[0039] (2)First iteration: input open source data and specific iteration topics as large model prompt words to generate the first chapter content;
[0040] (3)Subsequent iteration: input open source data, specific iteration topics and chapter content generated by the previous iteration as large model prompt words to generate subsequent chapter content;
[0041] (4)Integrate the answer: integrate the generated content in order to form a required article that overcomes the length limitation of the text.
[0042] In combination with the above technical solutions and the technical problems solved, the technical solution to be protected by the present application has the following advantages and positive effects:
[0043] First, the present application provides a large model assisted writing method based on knowledge enhancement and semantic coherence optimization, which has the following significant beneficial effects compared with the prior art:
[0044] (1)Real-time information collection and dynamic integration: build a content ecosystem with equal emphasis on timeliness and authority
[0045] One of the core innovations of the invention is the introduction of a real-time collection mechanism for multi-source heterogeneous data. By integrating data crawling interfaces for mainstream Internet platforms such as Twitter, Google, and WeChat public accounts, the system can instantly acquire and structure the latest global information. This mechanism breaks the limitations of traditional large models relying on static training data, enabling generated content to not only have the breadth of historical knowledge but also reflect current social hotspots, scientific research progress, or policy trends. The built-in dynamic information filtering engine has intelligent filtering algorithms that can automatically capture news, papers, policy documents, and other materials in the relevant field based on user-set theme keywords, and perform deduplication, abstract extraction, and credibility assessment. Time-sensitive content updates support periodic data source refreshing for time-sensitive reports, ensuring that the articles always reference the latest data.
[0046] (2) Personalized demand analysis: from "one face for thousands of people" to "one face for thousands of people" customized writing experience
[0047] The invention fully considers the writing style, language preference, and target audience characteristics of different users and proposes a personalized writing guidance mechanism based on prompt word engineering and style transfer. By customizing prompt word templates and style parameters, the system can accurately capture the user's expression intent and output text content that highly matches their expectations. Users can pre-set various prompt word combinations, such as "critical analysis," "objective statement," "encouraging conclusion," etc., and the system automatically adjusts the generation strategy based on the selected template. The style consistency maintenance module continuously monitors the language style (such as formality, sentence complexity, rhetorical devices) of the generated chapters during the multi-stage iterative generation process and maintains style consistency in subsequent generation to avoid abrupt transitions. User portrait modeling analyzes historical interaction data to gradually establish a user's writing style model, enabling automatic adaptation of the style without manual input.
[0048] For example, if a user wants to write a research report with academic rigor, the system will preferentially use passive voice, professional terminology, and long sentence structures; if the target reader is the general public, it will simplify the sentence structure and increase explanatory statements.
[0049] (3) Context coherence optimization: building a logically rigorous and progressive article architecture
[0050] To overcome the problems of "paragraph disconnection" and "logical jump" that are prone to occur when traditional large models generate long documents, the present application designs a large outline-driven staged iterative generation mechanism, which ensures high consistency of the entire article in terms of macro structure and micro semantic level through structured control and context memory reinforcement. By dividing the first-level title of the outline, multi-stage iteration is ensured to ensure the coherence of the article context. By expanding the context memory window, in each generation, the system not only provides the current task instruction, but also inputs the previously generated chapter content as context, helping the model understand the relationship between the front and back. Specifically, the generated outline is divided according to the first-level title, and each iteration not only includes open source data and specific iteration topics, but also includes the chapter content generated in the previous iteration as a reference, ensuring that the newly generated content is logically consistent with the existing content. For example, when writing a technical report, the first iteration generates the first chapter "Background Technology", and the subsequent iteration generates "Technical Solution", the large model will refer to the content of the first chapter to ensure that "Technical Solution" is logically consistent with "Background Technology", avoiding content jumps and semantic breaks. It can understand complex semantic structures and generate coherent text content, improving the overall quality of the article. It ensures that the logical relationship between chapters is close, and the entire article structure is rigorous and clear. The optimized article content is rich, improving the reader's reading experience, and meeting the requirements of professional report writing.
[0051] (4) Overcome the length limit of text generation: technical path for efficient support of long document generation
[0052] The present application proposes an innovative method to overcome the problem of difficult length control in a single generation of large models, which effectively supports the efficient generation of long professional reports through a combination of segmented generation and dynamic splicing, while considering generation quality and timeliness. In particular, to accurately control the overall length of the document, the present application introduces a dual control mechanism based on the number of chapters and the number of words per chapter.
[0053] In the outline design stage, the system will automatically generate a detailed outline structure according to the user input topic or topic. This outline contains multiple levels of titles, representing an independent writing unit. By analyzing the outline structure, it can be estimated how many such writing units the entire document will contain. This step provides a basis for subsequent word allocation. According to the user's requirements for the length of the final document, combined with the number of chapters estimated above, the system can calculate the ideal word range for each minimum chapter.
[0054] During the generation process, the system not only guides the content creation of each chapter according to the predetermined target word count, but also adjusts in real time according to the actual generation situation. At the same time, the system has a feedback loop to monitor the deviation between the actual word count of the generated chapter and the target value, and adjust the generation strategy of the remaining chapters accordingly, to ensure that the final document not only meets the expected length but also maintains the depth of content.
[0055] Secondly, through its unique technical means, including real-time information collection and dynamic integration, personalized demand analysis, context coherence optimization, and overcoming text generation length limitations, the invention not only significantly improves the efficiency and quality of report writing, but also significantly enhances user experience. In the business field, it can help enterprises quickly generate high-quality market analysis reports, product documents, etc., reducing labor costs and improving work efficiency. In education and academic research, it can help researchers quickly organize the latest research results and promote knowledge dissemination. In addition, for the news media industry, the invention can ensure that the article content keeps up with current events, improving the timeliness and accuracy of information to meet readers' demand for the latest information. In the long run, this not only promotes the informatization process of various industries, but also may give rise to new business models and service forms.
[0056] Although there are many writing assistance tools based on large models on the market, they generally have problems such as information update lag, insufficient personalized support, poor context coherence, and text generation length limitations. The invention effectively solves these challenges through innovative knowledge enhancement mechanisms, personalized cue word design, intelligent segmentation outline iteration, and multi-stage long text generation technology. In particular, it can dynamically obtain the latest data from multiple information sources and provide customized services according to user needs, making the generated articles both timely and in line with individual style preferences. This all-around solution fills many gaps in existing technology, especially in implementing long and complex document writing, reaching an unprecedented level.
[0057] For a long time, how to efficiently generate high-quality reports using large model technology has been a major problem for many scholars and practitioners. Especially in maintaining context coherence, breaking through text generation length limitations, providing personalized support, and ensuring information timeliness, existing methods are difficult to meet actual needs. Through in-depth research on these issues, the invention proposes a complete solution. For example, by dynamically searching and integrating the latest information resources, it ensures that the generated articles are closely related to the latest information; using personalized cue word design to meet users' specific needs and style preferences; using intelligent segmentation outline iteration to ensure the logical coherence of the article context; and overcoming text length limitations through multi-stage generation. These innovations collectively provide a feasible path to solve this long-standing technical problem.
[0058] The traditional view is that large models are more suitable for processing short text tasks and rely on static data sources, so they perform poorly in dealing with long and complex document writing. However, the present invention breaks this inherent concept and shows that through proper technical adjustment, large models can also play an important role in long writing. Specifically, the present invention emphasizes the search and integration of dynamic information resources and introduces personalized prompt word design, so that the system can flexibly adjust the output results according to the specific needs of the user. This method not only broadens the application range of large models, but also proves their adaptability and flexibility in different scenarios, thereby effectively overcoming the long-standing technical bias in the industry. BRIEF DESCRIPTION OF DRAWINGS
[0059] Figure 1 is a large model-based knowledge enhancement and context semantic coherent report writing generation method flowchart provided by an embodiment of the present invention;
[0060] Figure 2 is a large model-based knowledge enhancement and context semantic coherent report writing generation system architecture diagram provided by an embodiment of the present invention;
[0061] Figure 3 is a report quality evaluation comparison diagram of different models provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0062] In order to make the purpose, technical solutions and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below in combination with embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and do not limit the present invention.
[0063] The existing large models often have semantic isolation between the training data and the specific writing scene during report writing, resulting in generated text lacking pertinence and content depth. The present invention realizes semantic vectorization retrieval of the input theme by constructing a "knowledge enhancement module", and combines a Milvus vector database for high-dimensional cosine similarity matching and reordering. By filtering out the top 50 most relevant information to the theme, a highly correlated "dynamic knowledge subset" is established, significantly improving the knowledge fit of large models to target content and solving the information asymmetry problem between general corpus and specific writing needs of large models.
[0064] Traditional generated text tends to present a mechanical language style, lacking personalized expression and scene fit. To this end, the present system introduces a "personalized style writing topic and article outline module", which collects user preferences (such as professional term inclination, formality, expression style, etc.) and designs prompt word templates to drive large models to generate article topics and structural outlines that conform to user habits. This process is driven by semantic guidance and format specification, ensuring that the article theme is logically clear, structured, and has a distinct individualized style.
[0065] In existing text generation methods, there are often logical jumps or information duplication between chapters, which affects the structural rigor of long texts. To solve this problem, the system designs a "contextual semantic coherence optimization module" that uses a chapter-by-chapter progressive generation mechanism to divide the entire article into several first-level title paragraphs and dynamically integrates contextual relationships based on the generated content and the knowledge base double nested prompts. This ensures that each piece of content is independent and naturally connected in the overall semantic chain, avoiding content fragmentation and repetitive stacking from the root.
[0066] When large models generate long documents, they are often limited by the context length window, resulting in text truncation, information loss, or logical breaks. In this system, the "text generation length limitation overcoming module" solves this limitation through an iterative generation mechanism: the initial first chapter and knowledge base information are used to construct prompt words to generate the first paragraph of content, and then each iteration introduces "last chapter content + current title + background knowledge" to form a closed-loop prompt structure, effectively maintaining the continuous transmission of logical context and ensuring that the entire text is at a usable level in terms of structural continuity and semantic consistency.
[0067] The information sources introduced by the system include social media (such as Twitter), search engines (such as Google), and content platforms (such as WeChat public accounts) open multi-source data platforms. The heterogeneity and high-frequency update characteristics of this information ensure that the collected content is up-to-date and multi-perspective. Through unified vectorization processing and semantic clustering analysis, not only does it improve the contextual adaptability of background data, but also provides higher-dimensional reference semantic support for large models, making the generated report have both factual density and professional thickness.
[0068] The overall architecture of the system embodies a closed-loop writing mechanism in which multiple modules are prerequisites for each other and work together. Starting from knowledge-enhanced semantic adaptation, it moves on to personalized content generation control, then to recursive construction of segmented generation, structural constraints of semantic optimization, and finally to the integration of the entire text to achieve the output of a deliverable professional report. This system breaks away from the traditional "single generation-editing correction" logic chain and replaces it with a "knowledge-driven-structure decomposition-semantic tracking-content evolution" process system, completing the transition from AI-assisted writing to a professional task system solution.
[0069] As shown in Figure 1 The embodiment of the present application provides a knowledge-enhanced and contextually semantic coherent report writing generation method based on a large model, which specifically includes:
[0070] S1: The knowledge enhancement module receives user input topics or subjects, obtains relevant information from designated data sources according to the input topics, performs data collection, calculates the similarity between the topics and the information data in the database, and reorders the data to filter out the top 50 most relevant data, and stores the filtered data for use by the large model;
[0071] S2: The personalized style writing topic and article outline module collects user writing needs and style preferences, and the large model generates high-quality article topics according to user input topics or keywords, combined with user prompt word design and requirements, and generates standardized article outlines according to the topics and background knowledge base;
[0072] S3: The context semantic coherence optimization module divides the entire outline into multiple independent first-level titles according to the fine granularity of the first-level titles, ensuring that each chapter content can be processed independently. For each chapter, the large model prompt word technology is used to design and write content combined with the background knowledge base and the structure framework of the generated content, ensuring that each part is based on the latest information and consistent with other parts. If necessary, adjust to ensure that the entire article structure is rigorous and clear, and that each part of the content is related to each other, with a rigorous logic, improving the overall reading experience.
[0073] S4: The text generation length limitation overcoming module iteratively generates content segment by segment according to multiple outline sections. This method not only avoids the length limitation of single generation, but also improves the generation efficiency. The generated content of each section is integrated into a complete document according to the outline order to ensure that the final output document is complete in structure and clear in logic.
[0074] As shown in Figure 2 , the large model-based knowledge enhancement and context semantic coherence report writing generation system provided by the embodiment of the application specifically includes:
[0075] The knowledge enhancement module is used to provide the latest data preparation for report writing;
[0076] The personalized style writing topic and article outline module is used to generate article topics and standardized article outlines that meet the format requirements according to user input topics or keywords, and provide personalized prompt words to guide the large model to generate content;
[0077] The context semantic coherence optimization module is used to improve the logical coherence of long and complex report writing, making the generated report structure rigorous and clear, and meeting the requirements of professional report writing;
[0078] The text generation length limitation overcoming module is used to iteratively generate content segment by segment according to multiple outline sections.
[0079] The knowledge enhancement module is the preprocessing link of the entire invention method, and the specific implementation process is as follows:
[0080] (1)Information source selection: Select multiple information sources such as Twitter, Google, WeChat public number, etc. to provide free information search platform for all users;
[0081] (2)Information retrieval: Vectorize the collected data and store it in the vector database Milvus. According to the user's question, retrieve relevant background content from the selected information source;
[0082] (3)Similarity calculation: Vectorize the user's question and use cosine similarity to calculate the similarity of the retrieval results;
[0083] (4)Reordering: Reorder the retrieval results according to the similarity and select the top 50 most relevant results;
[0084] (5)Data integration: Integrate the top 50 most relevant results into open source data background knowledge for large model use.
[0085] The personalized style writing topic and article outline module is implemented as follows:
[0086] (1)Prompt word design: Design large model prompt words to guide the large model to generate article titles and outlines;
[0087] (2)Generate article title: Generate article title according to user input;
[0088] (3)Generate government document outline: Generate government document outline in strict accordance with markdown format.
[0089]
[0090]
[0091] The context semantic coherence optimization module uses the following code to divide the generated outline into multiple first-level titles according to the granularity of the first-level title.
[0092]
[0093] The design structure of the large model prompt word is as follows:
[0094]
[0095]
[0096] The length limitation of text generation module is implemented as follows:
[0097] (1)Segment generation: Generate content according to large model prompt word requirements.
[0098] (2) First iteration: input open-source data and specific iteration topics as large model prompts, generate the first chapter content. The structure is as follows:
[0099]
[0100]
[0101] (3) Subsequent iteration: input open-source data, specific iteration topics and chapter content generated by the previous iteration as large model prompts, generate subsequent chapter content. The structure is as follows:
[0102]
[0103] (4) Integration of answers: integrate the generated content in sequence to form a required article that overcomes the text length limit.
[0104] The specific implementation provided by the application is as follows:
[0105] (1) User input question: write a report on the recent 6.8 magnitude earthquake in Tibet. Data collection: the system obtains the latest reports and official statements about the 6.8 magnitude earthquake in Tibet from multiple sources such as Twitter, Google News, WeChat public account, China Earthquake Network, etc. Select the top 50 most relevant information to ensure that the background knowledge base contains the latest earthquake situation, rescue progress, expert analysis, etc. Store the filtered information for subsequent use of the large model.
[0106]
[0107] (2) Personalized style writing topic and article outline: the system identifies the user's requirement for a "rigorous" report and determines the need for formal and professional writing style to generate high-quality article topics and article outlines.
[0108]
[0109] (3) Intelligent segmentation of first-level title iteration: according to the first-level title of the outline, the generated task is divided into multiple sub-tasks of first-level title.
[0110]
[0111] (4) Segment generation: according to the outline, the document is divided into multiple parts, and the content is generated segment by segment. The generated content of each segment is integrated into a complete document according to the outline order to ensure that the final output document structure is complete and logical.
[0112] I. Earthquake overview
[0113] (I) Basic information about earthquake
[0114] 1. Earthquake occurrence time
[0115] On January 7, 2025, a magnitude 6.8 earthquake occurred in Tingli County, Xigazê City, Tibet Autonomous Region. The earthquake occurred at about 9:00 am local time, although it was short in duration, it was strongly felt, causing panic and evacuation among local residents. According to the China Earthquake Network, the earthquake occurred at about 9:00 on January 7, 2025, with the epicenter located in Tingli County, Xigazê City, Tibet Autonomous Region.
[0116] 2. Earthquake epicenter location
[0117] The present application provides a large model report writing method based on knowledge enhancement and semantic coherence optimization, aiming to dynamically search and integrate the latest information resources to ensure that the generated articles are closely related to the latest information; through personalized prompt word design, to meet the specific needs and style preferences of users; through intelligent segmentation outline iteration, to ensure the coherence of the context of the article; through multi-stage generation, to ensure that the length of the generated content breaks through the limit. The specific method is as follows:
[0118] Dynamic information resource integration and real-time updating technology: The present application dynamically obtains relevant information from Twitter, Google, WeChat public number and other major websites at home and abroad, then performs similarity calculation and reordering on the collected data, selects the top 50 most relevant information as the background knowledge base, and finally stores and manages the selected data for subsequent large model use. By integrating the latest information resources in real time, it ensures that the generated content keeps up with the latest information, provides the most advanced information support, and significantly improves the timeliness and accuracy of the article.
[0119] Personalized prompt word design and customized writing technology: Collect user writing needs and style preferences, conduct in-depth analysis of user needs, and large model prompt words simulate experienced writing experts to generate high-quality article titles and three-level title article outlines according to Markdown format. Provide highly personalized writing support to make the generated articles more in line with user expectations, significantly improve user experience and satisfaction.
[0120] Intelligent segmentation outline iteration and context coherence optimization technology: According to the first-level title of the outline, the generation task is divided into multiple title sub-tasks, and the background knowledge base and title content are reasonably prompted to ensure that the logical relationship between each part of the generated document is close, improving the overall quality of the article and making the reading experience more smooth and professional.
[0121] Multi-stage long text generation technology: divide the document into multiple parts according to the outline, generate content section by section, integrate the generated content of each section into a complete document in the order of the outline, and perform a final quality check on the entire document after integration to ensure seamless connection, no omissions or repetitions of the content of each part. Support efficient generation of long and complex report writing, break through the length limit of single generation, and significantly improve the generation efficiency and overall quality of writing.
[0122] It should be noted that the embodiments of the present application can be realized by hardware, software or a combination of software and hardware. The hardware part can be realized by special logic; the software part can be stored in the memory and executed by the appropriate instruction execution system, such as microprocessor or special design hardware. Those skilled in the art can understand that the above-mentioned devices and methods can be realized by computer executable instructions and / or included in processor control code, such as carrier medium, such as magnetic disk, CD or DVD-ROM, programmable memory, such as read-only memory (firmware), or data carrier, such as optical or electronic signal carrier. The device and its modules of the present application can be realized by hardware circuit, such as ultra-large scale integrated circuit or gate array, semiconductor, such as logic chip, transistor, etc., or programmable hardware device, such as field programmable gate array, programmable logic device, etc., can also be realized by software executed by various types of processors, and can also be realized by the combination of the above hardware circuit and software, such as firmware.
[0123] The above is only a specific embodiment of the present application, but the protection scope of the present application is not limited thereto. Any modification, equivalent replacement and improvement made by those skilled in the art within the technical range disclosed by the present application, as long as it is within the spirit and principle of the present application, should be covered within the protection scope of the present application.
Claims
1. A large model-based knowledge enhancement and context semantic coherence report writing generation method, characterized in that, The method comprises the following steps: receiving user input topic information, extracting text data related to the topic from preset multi-source data using vectorization retrieval, calculating similarity and sorting according to relevance, and selecting the top 50 texts with the highest relevance to construct a background knowledge set; constructing prompt words according to the user's writing style parameters, generating an article title and a standardized article outline; dividing the target text into chapters according to the article outline, generating chapter text for each chapter by driving the large model under the input of the background knowledge set and the generated content, and controlling semantic coherence between chapters; generating all chapter texts in turn using an iterative generation mechanism, and integrating the complete report in chapter order.
2. A large model-based knowledge enhancement and contextually semantically coherent report writing generation system, characterized by, It includes: a knowledge enhancement module for retrieving texts related to user input topics from multi-source data and filtering the top 50 texts with the highest relevance to construct a background knowledge set; a style and structure generation module for collecting user writing style parameters and generating an article title and a standardized article outline; a context control module for injecting the background knowledge set and generated content when generating chapter texts to maintain semantic coherence; a segmented generation module for generating chapter texts in turn through an iterative generation mechanism and completing integration.
3. An electronic device comprising a processor and a memory, the memory storing computer executable instructions, when the instructions are executed by the processor, causing the electronic device to perform the method of claim 1.
4. A computer readable storage medium storing computer executable instructions, the instructions, when executed by a processor, causing the processor to perform the method of claim 1.
5. The system of claim 2, wherein, The knowledge enhancement module includes an information retrieval unit, a similarity calculation unit, a reordering unit and a knowledge integration unit, and the knowledge integration unit is used to convert the filtering result into a structured background knowledge set.
6. The system of claim 2, wherein, The style and structure generation module includes a user parameter acquisition unit, a prompt word construction unit and an outline generation unit, and the outline generation unit is used to output a hierarchical structure in conformity with the tokenization format.
7. The system of claim 2, wherein, The context control module establishes a chapter dependency graph based on the article outline, and injects the structure summary and keyword information of the previous chapter when generating the current chapter text.
8. The system of claim 2, wherein, The segmented generation module inputs the last chapter text, the chapter title and the background knowledge set to construct new prompt words when generating subsequent chapters to drive the large model to generate.
9. The method of claim 1 wherein, Before one step, there is also a preprocessing step of removing theme irrelevant data and performing text deduplication to improve the accuracy of the background knowledge set.
10. The method of claim 1, wherein, The multi-source data includes social media text, search engine public web page text and content publishing platform text.
Citation Information
Cited By
Report generation method and system based on large language model and multi-source information fusion
CN121188091A
Report generation method and system based on large language model and multi-source information fusion
CN121188091B
Knowledge-enhanced multi-dimensional controllable text generation method and system
CN121328747A
Report generation method and device and computer program product
CN121388195A