Report generation method and system based on large language model and multi-source information fusion

By using a large language model and multi-source information fusion, a hierarchical report outline is generated and dynamic information is integrated, which solves the problems of low report generation efficiency, knowledge lag, and lack of data support, and achieves efficient, professional and accurate report generation.

CN121188091BActive Publication Date: 2026-02-06UESTC (SHENZHEN) ADVANCED RES INST
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511739140.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-11-25
Publication Date
2026-02-06
Estimated Expiration
2045-11-25

AI Technical Summary

Technical Problem

Existing report generation methods are inefficient, prone to "illusions," suffer from outdated knowledge, lack data support, and cannot adapt to complex and ever-changing analytical needs.

Method used

The method adopts a large language model and multi-source information fusion approach. It generates a hierarchical report outline, dynamically integrates information to generate chapter text, and utilizes professional document knowledge bases, online information sources and industry databases to support multi-source information fusion and automatically create citation tags.

Benefits of technology

It has achieved a fully automated process from understanding requirements to formatting, ensuring the professionalism, timeliness and accuracy of reports, reducing the 'illusion' phenomenon, and improving the efficiency and rigor of report generation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121188091B_ABST
    Figure CN121188091B_ABST
Patent Text Reader

Abstract

The application discloses a report generation method and system based on a large language model and multi-source information fusion, comprising: generating a report outline containing hierarchical chapters based on user input and a predefined structured template; for the chapters in the report outline, generating chapter texts through an iterative process, including: generating key questions based on the content of the chapters by a domain large model, performing dynamic information fusion processing on each key question to generate professional answers, and synthesizing chapter texts based on all professional answers in the chapters; combining and formatting all chapter texts according to the report outline to output a complete report. By generating key questions for each chapter, the chapter target is converted into a series of specific and answerable subtasks; through multi-source information fusion, the report content is based on an authoritative professional knowledge base and is jointly supported by the latest current news and accurate industry data, greatly reducing the "illusion" phenomenon and ensuring the professionalism, timeliness and accuracy of the content.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of natural language processing, in particular to a report generation method and system based on a large language model and multi-source information fusion. BACKGROUND

[0002] With the development of information technology, the demand for professional, in-depth, and timely industry reports in the decision-making process of various industries is increasing. The traditional report writing method relies on manual work, which usually includes data collection, data query, literature reading, content writing, and format layout. This method has the problems of low efficiency, high cost, long cycle, and being easily limited by the professional level of personnel. Some automatic document generation schemes have appeared in the prior art, which generate reports through pre-defined fixed templates and filling rules. The content is rigid and stereotyped, and cannot adapt to complex and variable analysis requirements. Using a large language model to directly generate text according to user instructions relies on the parameterized knowledge of the model itself, which is prone to "hallucination" problems and cannot cover the latest market trends, policy changes, and news events, resulting in knowledge lag. The generated qualitative discussion often lacks precise quantitative data as support, resulting in insufficient report persuasiveness. SUMMARY

[0003] The purpose of the present application is to provide a report generation method and system based on a large language model and multi-source information fusion to solve the problems of low report generation efficiency, "hallucination", knowledge lag, and lack of data support mentioned in the background art.

[0004] To achieve the above purpose, the technical scheme adopted by the present application is as follows:

[0005] According to one aspect of the present application, a report generation method based on a large language model and multi-source information fusion is provided, which comprises:

[0006] Based on user input and a pre-defined structured template, a report outline containing hierarchical chapters is generated;

[0007] For the chapters in the report outline, the chapter text is generated through an iterative process, including: generating key questions based on the content of the chapter by a domain large model, performing dynamic information fusion processing to generate professional answers for each key question, and synthesizing the chapter text based on all professional answers in the chapter;

[0008] All chapter texts are combined and formatted according to the report outline to output a complete report.

[0009] Based on the foregoing scheme, the dynamic information fusion processing is performed on each key question to generate a professional answer, including: generating an initial answer based on the retrieval result of the professional document knowledge base; analyzing the initial answer and dynamically triggering the acquisition of time-sensitive supporting information from a network information source and data support information from an industry database; and performing multi-source information fusion on the initial answer, the time-sensitive supporting information, and the data support information to generate the professional answer.

[0010] Based on the foregoing scheme, the initial answer is generated based on the retrieval result of the professional document knowledge base, including: performing hybrid retrieval of semantic retrieval and keyword retrieval on the key question, reordering and fusing the retrieval result, and generating the initial answer based on the preferred retrieval material after reordering and fusion by the domain large model.

[0011] Based on the foregoing scheme, the reordering and fusion adopts a reverse ranking fusion algorithm, calculates a fusion score based on the ranking of document slices in the semantic retrieval and keyword retrieval results, and selects the top N document slices with the highest scores as the preferred retrieval material.

[0012] Based on the foregoing scheme, the time-sensitive supporting information is acquired from a network information source, including: generating a network retrieval instruction containing a time range and a core topic by the domain large model, and after obtaining preliminary results through a network search engine interface, performing risk filtering on the information by combining keyword matching and security model reasoning.

[0013] Based on the foregoing scheme, the data support information is acquired from an industry database, including: generating an industry data query instruction by the domain large model, converting it into a database query statement through Text-to-SQL technology, and summarizing the results into text description of industry data material by the domain large model after querying.

[0014] Based on the foregoing scheme, the multi-source information fusion includes: taking the initial answer as a discussion skeleton, embedding the time-sensitive supporting information as empirical cases, and integrating the data support information as core arguments, and generating the professional answer by the domain large model through semantic understanding and text rewriting.

[0015] Based on the foregoing scheme, in the process of multi-source information fusion, reference identifiers corresponding to the cited information sources are automatically created in the generated text.

[0016] Based on the foregoing scheme, the generation of a report outline containing hierarchical chapters includes: extracting key attributes from the user input and filling in reserved spaces of the structured template to generate a machine-readable report outline data object.

[0017] In accordance with another aspect of the present application, there is provided a report generation system based on large language models and multi-source information fusion, comprising a report generation engine, a professional document knowledge base, an information acquisition and filtering module, an industry database, a report template library;

[0018] The report generation engine is configured to perform operations of report outline generation, chapter text generation, and report output.

[0019] The professional document knowledge base is configured to provide search results based on professional documents to the report generation engine.

[0020] The information acquisition and filtering module is configured to provide risk-filtered current affairs materials to the report generation engine.

[0021] The industry database is configured to provide data materials to the report generation engine.

[0022] The report template library is configured to provide predefined structured templates to the report generation engine.

[0023] According to the above technical solution, the present application has at least the following advantages and positive effects compared with the prior art: by generating key questions for each chapter, the chapter target is converted into a series of specific and answerable sub-tasks; through multi-source information fusion, the report content is based on an authoritative professional knowledge base and is supported by the latest current affairs news and accurate industry data, greatly reducing the "illusion" phenomenon and ensuring the professionalism, timeliness and accuracy of the content. The present application realizes an end-to-end automatic process from demand understanding, outline construction, content generation to format typesetting; simulates expert thinking, dynamically analyzes information gaps and intelligently triggers supplementary retrieval, significantly reduces manual intervention and improves report generation efficiency. By introducing a predefined information source priority strategy, potential contradictions between different information sources can be automatically detected and arbitrated to ensure that the final argument is based on the most reliable information, enhancing the rigor and authority of the report.

[0024] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present application. BRIEF DESCRIPTION OF DRAWINGS

[0025] The accompanying drawings, which are incorporated into and form part of the specification, illustrate embodiments consistent with the present application and, together with the specification, serve to explain the principles of the present application. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained from these drawings without creative labor for those skilled in the art. In the drawings:

[0026] Figure 1A flowchart of the report generation method based on a large language model and multi-source information fusion of the present application is shown in Figure 1.

[0027] Figure 2 A flowchart of the report generation method based on a large language model and multi-source information fusion of the present application is shown in Figure 1. DETAILED DESCRIPTION

[0028] In order to more clearly illustrate the purpose, technical scheme and advantages of the present application, the technical scheme in the embodiments of the present application will be described in detail below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, not all the embodiments. The example embodiments can be implemented in various forms, and should not be understood as being limited to the examples described herein. On the contrary, these embodiments are provided to make the present application more comprehensive and complete, and to fully convey the ideas of the example embodiments to those skilled in the art.

[0029] In addition, the described features, structures or characteristics can be combined in any suitable manner in one or more embodiments. In the following description, many specific details are provided to give a sufficient understanding of the embodiments of the present application. However, those skilled in the art will realize that the technical solutions of the present application can be practiced without one or more of the specific details, or other methods, components, devices, steps, etc. can be used. In other cases, well-known methods, devices, implementations or operations are not shown or described in detail to avoid obscuring the aspects of the present application.

[0030] The block diagram shown in the accompanying drawings is only a functional entity, which does not necessarily correspond to a physically independent entity. That is, these functional entities can be implemented in the form of software, or in one or more hardware modules or integrated circuits, or in different networks and / or processor devices and / or microcontroller devices.

[0031] The flowchart shown in the accompanying drawings is only an exemplary illustration, which does not necessarily include all the contents and operations / steps, and is not necessarily executed in the order described. For example, some operations / steps can be further decomposed, and some operations / steps can be combined or partially combined, so the actual execution order can be changed according to the actual situation.

[0032] The present application will be described in detail below with reference to specific embodiments.

[0033] Embodiment 1

[0034] As shown in Figure 1 , 2 The present embodiment provides a report generation method based on a large language model and multi-source information fusion, and the specific steps of the method are as follows:

[0035] S1: Based on user input and predefined structured templates, generate a report outline containing hierarchical chapters.

[0036] Based on user input, extract key attributes through intelligent parsing, and instantiate using predefined structured templates to generate a machine-readable report outline containing hierarchical chapters.

[0037] Intelligent parsing and key information extraction based on user needs; receive user input, which includes free-form natural language text, such as: "Help me write a report on the application prospects of artificial intelligence in the field of smart medical care"; or specific fields provided through structured forms.

[0038] For free text input, call the domain-adaptive language model fine-tuned by instructions for deep semantic understanding of the input text; the model accurately extracts core entities and key attributes related to the report topic from the user description by performing sequence-to-sequence conversion and information extraction tasks in its encoder-decoder architecture; these attributes include at least: report core topic, industry category, professional keywords, possible time range, geographical range, etc. For structured input, use rule matching and field mapping mechanism, directly capture and standardize field content through predefined parsing rules. Preferably, if the key attributes extracted from the free text are missing or ambiguous, an interactive clarification round can be initiated, or relevant attributes can be supplemented according to the predefined default strategy to ensure the completeness of the information.

[0039] Further, construct a report outline based on a structured template, predefine a professional report template library, where each template is a structured data object defined in JSON format, which hierarchically defines the title, chapter title, chapter level and content description of the report. The standardized key attributes extracted above (such as core topic, industry category) are accurately filled into the corresponding structured field vacancies (such as "report_topic", "industry", "key_terms") in the JSON template through the template instantiation process. After filling, a structured, machine-readable report outline data object is generated; this data object serves as the basis for driving subsequent chapter content generation, clearly specifying the final title of the report, the sequence of chapters to be generated, the hierarchical depth of each chapter, and giving each chapter a content-oriented context with key attributes.

[0040] S2: Generate chapter text for the chapters in the report outline through an iterative process.

[0041] S201: Generate key questions based on the content of the chapter by the domain model.

[0042] The process iterates through each chapter in the report outline data object. First, the domain big data model intelligently analyzes the core theme of each chapter. Based on the chapter's hierarchy and the complexity of its content description, the domain big data model simulates expert thinking to deconstruct the chapter's theme. Based on the question generation pattern learned during the instruction fine-tuning phase, the model autonomously generates multiple key professional questions (critical questions) by analyzing the chapter title, its hierarchy within the report, the content description from the template, and the core keywords already filled in during template instantiation. Higher-level, overview chapters (such as the introduction) correspond to fewer questions (e.g., 1-2), while lower-level, in-depth analytical chapters (such as technical detail analysis) require specific arguments and thus correspond to more (e.g., 3-5) micro-level questions. When generating each critical question, the model uses a self-attention mechanism in its decoder to ensure that the question semantics are aligned with the chapter's core theme. Simultaneously, the model employs diversity constraint techniques in sequence generation to ensure that the question group systematically covers all core aspects of the chapter (e.g., definition, current status, trends, challenges) and minimizes semantic overlap between questions. Among them, diversity constraint technology refers to the computational strategy adopted in the language model decoding and generation process to reduce the repetition of generated text and increase content diversity. Its implementation methods include, but are not limited to: penalizing generated lexical units to reduce their probability of recurrence in new questions; adopting a kernel sampling (Top-p sampling) strategy to randomly select from a set of high-probability candidate words to introduce diversity; introducing diversity measurement in beam search to promote semantic dispersion of multiple generated questions; through the combination of one or more of the above techniques, the model can automatically generate a set of key questions that are both related and mutually exclusive.

[0043] Ultimately, all generated questions are designed to be responsive to at least one information source from subsequent professional document knowledge bases, online search engines, or industry databases, thus laying a precise query framework for subsequent information retrieval and integration. For example, for the "Market Competitiveness Analysis" chapter, questions such as "What is the market concentration of this industry?" and "What are the core advantages of the main competitors?" might be generated, transforming the chapter's objectives into a series of specific, answerable sub-tasks.

[0044] S202: Perform dynamic information fusion processing on each of the key questions to generate a professional answer.

[0045] Each generated key question is traversed, and a multi-stage, dynamic information retrieval and fusion process is executed to generate a final professional answer. An initial answer is generated based on the search results from a professional document knowledge base; the initial answer is analyzed, and timely supporting information is dynamically retrieved from online information sources and structured data support information is retrieved from industry databases; the initial answer is then fused with the timely supporting information and structured data support information to generate a professional answer.

[0046] Firstly, the key question is submitted to the professional document knowledge base; the retrieval results (such as relevant document snippets) are obtained through a hybrid retrieval mechanism, and the retrieval results are fused and selected using a reordering model to obtain the preferred retrieval materials; the domain large model generates an initial answer based on these preferred retrieval materials; this answer constitutes the basis for the professional and qualitative discussion of the key question.

[0047] Among them, the hybrid retrieval mechanism executes two retrieval paths in parallel: semantic retrieval based on dense vectors and keyword retrieval based on keyword vectors, to balance semantic relevance and keyword matching accuracy. For the semantic retrieval path, a pre-trained embedding model (such as bge-m3) is used to encode the key question into a high-dimensional dense vector; this vector captures the deep semantic information of the question in the vector space; an approximate nearest neighbor search is performed in the Milvus vector database to find the top K document snippets with the highest cosine similarity to the question vector; the semantic retrieval path can effectively identify document content that is semantically related to the question but does not overlap in keywords. For the keyword retrieval path, core keywords are extracted from the key question, and a sparse vector is generated using the embedding model or a traditional keyword inverted index is used for full-text retrieval in the database; the keyword retrieval path can accurately match documents containing specific professional terms and ensure that core concepts are not missed.

[0048] After obtaining the preliminary retrieval results of the two paths, a reordering model is called to fine-tune the merged result list; the reordering model uses a reverse ranking fusion algorithm to assign scores to each item's ranking in different lists and perform weighted calculations to obtain a final fusion score. Specifically, a document snippet ranks R vec in the semantic retrieval results and ranks R kw in the keyword retrieval results, then its RRF score is: ; where k is the damping constant, usually taken as 60; select the top N (such as 5) document snippets with the highest final score as the preferred retrieval materials passed to the large model.

[0049] Further, the key question and the preferred retrieval materials are combined into a continuous text sequence as input to the domain large model; the input sequence follows a specific instruction template, for example, the instruction template can be "Answer the question based on the following knowledge: [preferred retrieval material 1]... [preferred retrieval material N]; question: {key question}". The model's tokenizer converts this input sequence into a word token sequence, and the model's embedding layer converts each word token into a high-dimensional vector representation; this vector not only contains the semantic information of the word token, but also indicates its order in the sequence through position encoding, thereby laying the foundation for the model to understand the context structure.

[0050] Within the model, the retrieved materials are subjected to deep understanding and correlation analysis through its attention mechanism, identifying the most relevant evidence and arguments to the question. The encoded input sequence enters the model's Transformer decoder architecture, activating multi-layer attention and cross-attention mechanisms to perform key information processing tasks. At the bottom layer of the model, the self-attention mechanism enables each token in the input sequence to interact globally with all other tokens in the sequence. This allows the model to fully understand the logical relationships between sentences and arguments within the preferred retrieval materials; accurately grasp the semantics of each word in the key question and their grammatical structure. At a deeper level of the model, in the cross-attention mechanism, the vector representing the key question is used as a query to actively examine all vectors representing the preferred retrieval materials (as keys and values); by calculating attention weights, the model can identify the most relevant fragments, data, and arguments in the materials to the question; high-attention-weighted material fragments will be given higher prominence and converge into an information-rich context vector, serving as a direct basis for generating answers; thus achieving precise positioning of evidence from massive retrieval information and effectively suppressing irrelevant information interference.

[0051] After fully understanding the question and locking in the key evidence, the model enters the answer generation phase, i.e., constructing an initial answer. The model is strictly instructed to use the provided retrieval materials as factual basis and knowledge boundary, and the model's generated content is constrained within the scope of the retrieved evidence, greatly reducing the possibility of producing "hallucinations" or fabricating facts. The model generates the initial answer in a self-recursive manner, one token at a time; specifically, the model calculates the probability distribution of the first token based on the context of the entire input sequence through the Softmax output layer, and selects the most likely token (e.g., "artificial intelligence"); the generated token is appended to the end of the input sequence, forming a new sequence; the model, based on the new sequence, again predicts the next token through its attention mechanism; at this time, the causal mask is used to ensure that only the generated content can be focused on. This process is iterated until a complete text sequence, i.e., the initial answer, is generated.

[0052] The domain large model then conducts a critical analysis of the initial answer, including classification of the nature of the arguments in the text to identify information gaps, and dynamically triggers two types of supplementary retrieval: classifying arguments related to trends, policy directions, and recent events as time-sensitive arguments (requiring timely information), and constructing web search instructions accordingly; classifying arguments related to market size, share, growth rate, and statistical ranking as data-supported arguments (requiring data support), and constructing industry data query instructions accordingly.

[0053] For time-sensitive arguments, the large model analyzes the initial answer, identifies points in the argument that require the latest news, policies, or market dynamics for support or reinforcement, and automatically constructs precise web search instructions that explicitly include two core fields: the time range of the query (such as time_range) and the core topic (such as core_topic). After obtaining preliminary results through a web search engine interface, the information is cleaned through keyword matching (first-level filtering) and a security review model (second-level filtering), and finally the safe and relevant current event materials (time-sensitive supporting information) are output. Among them, keyword matching quickly matches the preliminary results obtained with a predefined sensitive word database, marking or removing entries containing high-risk words; through the security review model (such as qwen3-guard), the first filtered information is subjected to deep semantic understanding, and risk content such as implicit bias, discrimination, and privacy leakage is identified and filtered out, finally outputting clean and safe current event materials.

[0054] For data-supported arguments, the large model simultaneously analyzes the initial answer and identifies arguments that require specific numerical values, market size, growth rate, and other quantitative data support. For these arguments, the large model automatically constructs an industry data query instruction. Through Text-to-SQL technology, the instruction is converted into an executable database query statement, the query is completed in the industry database, and the large model summarizes the queried structured data into a text description of industry data materials.

[0055] Further, the multi-source information is intelligently fused and the final answer is generated. The domain large model takes the initial answer as the discussion skeleton, fuses the current event materials and industry data materials, embeds the time-sensitive supporting information (current event materials) as empirical cases, integrates the data support information (industry data materials) as core arguments, and generates a professional answer through semantic understanding and text rewriting by the domain large model. Before fusion, information consistency verification is first performed. The model compares the statements and data from different information sources through its attention mechanism, identifies potential logical contradictions or factual conflicts, for example, the historical market size in the knowledge base does not match the latest data in the database. The model resolves potential conflicts between different information sources according to a predefined and configurable information source priority strategy to ensure the authority and accuracy of the final answer; for example, the priority order can be: accurate data from industry databases > high timeliness information obtained from network search engines > professional document knowledge base documents. When a conflict is detected, the model will automatically adopt the core facts and data from the higher priority information source, and use it as a reference to modify the facts and update the data in the corresponding part of the discussion skeleton. The large model performs semantic understanding and text rewriting, dynamically embeds the relevant context of the argument in the form of quotation or bracket supplement, and provides time-sensitive supporting evidence for it. The industry data materials are used as core arguments to quantify and strengthen the qualitative description through data assimilation and creation of data support sentences, etc., to generate a logically coherent final professional answer. For example, "rapid growth" is rewritten as "achieved XX% significant growth". Preferably, in this process, according to the cited information source, a corresponding reference identifier is automatically created in the generated text to realize full-process information tracing.

[0056] S203: Synthesize the chapter text based on all the professional answers within the chapter.

[0057] When all the key questions of a chapter have obtained the corresponding final professional answers, the domain large model takes these answers as core materials and combines them with the already generated preface report content to perform global context connection. The large model is responsible for eliminating content duplication, ensuring logical progression, adding transition sentences, and finally synthesizing a coherent, complete, and academically compliant chapter text.

[0058] Through the recursive process of question generation-answer construction-text synthesis, and by introducing a static-dynamic intelligent information fusion mechanism in answer construction, the method triggers retrieval on demand based on dynamic analysis and organically integrates information of different natures.

[0059] S3: Combine and format all the chapter texts according to the report outline and output a complete report.

[0060] The chapter text contents generated in step S2 are spliced in the order and hierarchy specified by the report outline data object generated in step S1. Specifically, a global document object model is maintained in memory, and each chapter is inserted as an independent logical unit into its predefined position in the report outline, thereby completing the content assembly of the complete report. After the content assembly is completed, the entire text is scanned to perform a term consistency check, ensure uniformity of core concept expression, and eliminate content duplication or logical gaps that may be caused by chapter generation. A document rendering engine is called to automatically apply professional layout styles to different levels of headings, text, charts, and other elements according to predefined document style rules. Preferably, the title structure in the document can be automatically parsed, and a report table of contents can be dynamically generated and inserted. Maintainable cross-references for charts and chapters can be created, and the system can automatically generate and maintain their numbering and hyperlinks to ensure convenient and accurate internal navigation of the document. If reference identifiers (superscripts) are generated in step S2, all reference information will be collected at this stage, and a unified list of references or data source explanations will be automatically generated at the end of the report to complete the information traceability loop.

[0061] After all integration and formatting is completed, the document object model in memory is serialized into one or more standardized file formats for output. These formats include but are not limited to: editable document formats (such as.docx), which preserve all styles and structures for user fine-tuning; fixed layout formats (such as.pdf), which ensure consistent display on any device and are suitable for formal publication and circulation; and structured data formats (such as.json), which contain complete text content and metadata for further processing and analysis by other systems or databases. Users can choose one or more formats to receive the final industry report according to their needs. Preferably, report version identifiers and data source snapshot lists can be generated to ensure traceability. Further, multi-modal publishing through API is supported to push the report or its summary to designated content management systems or automatically generate a presentation.

[0062] Embodiment 2

[0063] This embodiment exemplarily presents a report generation system based on large language models and multi-source information fusion, including a report generation engine, a professional document knowledge base, an information acquisition and filtering module, an industry database, and a report template library.

[0064] The report generation engine is used to perform report outline generation, chapter text generation, and report output operations; receives user requests, calls the report template library to generate an outline, and drives the entire process from chapter question generation, multi-source information retrieval and fusion, to final report synthesis and formatting; wherein the built-in domain large model is the core of intelligent analysis and content generation.

[0065] A professional document knowledge base is configured to provide search results based on professional documents to the report generation engine. The knowledge base is built based on a vector database (e.g., Milvus) and stores cleaned and structured industry documents (e.g., research reports and academic papers). The knowledge base provides hybrid search capabilities and can quickly locate the most relevant search materials from massive documents through semantic and keyword matching.

[0066] An information acquisition and filtering module is configured to provide risk-filtered current event materials to the report generation engine. The module acquires network information (e.g., news and policies) with high timeliness by calling a network search engine interface. The module automatically identifies and filters sensitive and illegal content by combining keywords and a security model, thereby ensuring the cleanliness and safety of the input information.

[0067] An industry database is configured to provide data materials to the report generation engine. The database stores structured industry statistical data (e.g., market size and enterprise revenue) and supports the conversion of natural language questions into database query statements through Text-to-SQL technology. The database can automatically summarize the query results into a text description, thereby providing quantitative support for the report.

[0068] A report template library is configured to provide predefined structured templates to the report generation engine. The library stores various professional report format templates in a structured format (e.g., JSON). The templates define the report title, chapter structure, hierarchical relationship, and content description. The templates reserve spaces to realize the rapid initialization and structured generation of the report, thereby ensuring the standardization and professionalism of the output content.

[0069] Other embodiments of the application will be apparent to those skilled in the art from consideration of the specification and practice of the application disclosed herein. It is intended that the specification and examples be considered as exemplary only, with the true scope and spirit of the application being indicated by the following claims. It is to be understood that the application is not to be limited to the precise details of construction and operation illustrated and described herein, and that various modifications, changes and substitutions in the specifics thereof can be made by those skilled in the art without departing from the scope and spirit of the application. The scope of the application is limited solely by the claims that follow.

Claims

1. A report generation method based on a large language model and multi-source information fusion, characterized in that, The method comprises: generating a report outline containing hierarchical chapters based on user input and a predefined structured template; generating chapter texts for chapters in the report outline through an iterative process, including: generating key questions by a domain large model based on the content of the chapters, performing dynamic information fusion processing on each key question to generate professional answers, and synthesizing chapter texts based on all professional answers within the chapters; the dynamic information fusion processing includes: generating an initial answer based on the retrieval results of a professional document knowledge base; analyzing the initial answer and dynamically triggering the acquisition of time-sensitive supporting information from network information sources and data support information from industry databases; and performing multi-source information fusion on the initial answer, the time-sensitive supporting information, and the data support information to generate the professional answer; combining and formatting all chapter texts according to the report outline to output a complete report.

2. The report generation method based on a large language model and multi-source information fusion according to claim 1, characterized in that, The generation of the initial answer based on the retrieval results of the professional document knowledge base includes: hybrid retrieval of semantic retrieval and keyword retrieval on the key questions, reordering and fusion of the retrieval results, and generation of the initial answer by the domain large model based on the preferred retrieval materials after reordering and fusion.

3. The report generation method based on a large language model and multi-source information fusion according to claim 2, characterized in that, The reordering and fusion adopts a reverse ranking fusion algorithm, calculates a fusion score based on the ranking of document slices in the semantic retrieval and keyword retrieval results, and selects the top N document slices with the highest scores as the preferred retrieval materials.

4. The report generation method based on a large language model and multi-source information fusion according to claim 1, characterized in that, The acquisition of time-sensitive supporting information from network information sources includes: generating a network retrieval instruction containing a time range and a core topic by the domain large model, and performing risk filtering on the information through keyword matching and security model reasoning after obtaining preliminary results through a network search engine interface.

5. The report generation method based on a large language model and multi-source information fusion according to claim 1, characterized in that, The acquisition of data support information from industry databases includes: generating an industry data query instruction by the domain large model, converting it into a database query statement through Text-to-SQL technology, and summarizing the results into text description of industry data materials by the domain large model after querying.

6. The report generation method based on a large language model and multi-source information fusion according to claim 1, characterized in that, The multi-source information fusion includes: taking the initial answer as the discussion framework, embedding the time-sensitive supporting information as empirical cases, and integrating the data support information as core arguments, and generating the professional answer by the domain large model through semantic understanding and text rewriting.

7. The report generation method based on a large language model and multi-source information fusion according to claim 6, characterized in that, In the process of multi-source information fusion, reference identifiers corresponding to the cited information sources are automatically created in the generated text.

8. The report generation method based on a large language model and multi-source information fusion according to claim 1, characterized in that, The generation of the report outline containing hierarchical chapters includes: extracting key attributes from the user input and filling in the reserved spaces of the structured template to generate a machine-readable report outline data object.

9. The report generation method system based on large language model and multi-source information fusion, for implementing the report generation method according to any one of claims 1-8, characterized in that, It includes a report generation engine, a professional document knowledge base, an information acquisition and filtering module, an industry database, and a report template library; The report generation engine is used to perform report outline generation, chapter text generation, and report output operations; The professional document knowledge base is used to provide retrieval results based on professional documents to the report generation engine; The information acquisition and filtering module is configured to provide risk-filtered news material to the report generation engine. The industry database is configured to provide data material to the report generation engine. The report template library is configured to provide predefined structured templates to the report generation engine.

Citation Information

Patent Citations

  • Long report generation method based on large model

    CN120471016A

  • Knowledge enhancement and context semantic coherence report writing generation method and system based on large model

    CN120911587A