A report generation method combining artificial intelligence and multi-path vector recall
Patent Information
- Application Number
- CN202610634278.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-09
- Publication Date
- 2026-09-11
AI Technical Summary
内容原创度不足:大量依赖公开网络资源,缺乏深度信息加工与语义重构,导致生成内容与实际业务需求的契合度偏低,原创性和参考价值不足
[0054]与现有技术路径相比,本发明在多方面形成了显著的技术优势与应用价值。
Smart Images

Figure CN122735641A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of natural language processing and information retrieval technology, and in particular to a method and system for rapid generation of research reports based on text embedding models, vector databases, multi-path semantic retrieval fusion, and generative large language models. Background Technology
[0002] Currently, the mainstream technical approach for automatically generating research reports largely relies on web scraping combined with template-based content assembly. These systems have the following significant limitations in practical applications: Insufficient originality of content: It relies heavily on public online resources and lacks in-depth information processing and semantic reconstruction, resulting in a low degree of alignment between the generated content and actual business needs, and insufficient originality and reference value.
[0003] Limited customization capabilities: It is difficult to generate highly personalized, structured, and relevant texts to meet the specific needs of users, and the generated reports often have deficiencies in terms of language fluency and logical structure.
[0004] Limited recall coverage: Most solutions rely on a single knowledge source or a single retrieval strategy, resulting in unstable comprehensiveness, timeliness, and semantic relevance of information recall.
[0005] Poor generation quality: Lacking advanced natural language processing techniques (such as retrieval enhancement generation and context optimization), the generated results are deficient in terms of logical coherence, timeliness support, and accuracy of details.
[0006] These shortcomings directly impact enterprises' ability to obtain accurate, real-time, and customizable reports in scenarios such as market research and strategic decision-making. Therefore, a high-performance technology system is needed that integrates multi-source knowledge, supports multi-channel semantic retrieval, and incorporates generative AI reasoning capabilities to achieve rapid and high-quality production of research reports. Summary of the Invention
[0007] To address the problems existing in the prior art, the purpose of this invention is to provide a report generation method that combines artificial intelligence and multi-path vector recall. This invention integrates multi-source knowledge retrieval, vector database indexing, multi-path semantic recall, and generative artificial intelligence reasoning technologies, enabling the rapid generation of high-performance research reports. It overcomes the technical bottlenecks of existing automated research report generation systems in terms of content originality, customized structure, comprehensive information recall, and logical rigor.
[0008] By introducing a multi-vector recall mechanism and a Retrieval-Augmented Generation (RAG) strategy, this invention can improve the business fit and structured quality of text generation while ensuring the relevance and coverage of the recalled data, thereby achieving a high-precision, scalable, and traceable report generation process for different industry sectors.
[0009] Existing automated research report generation technologies have revealed several constraints in practical business implementation, including but not limited to the following aspects.
[0010] 1. Lack of high-quality original content and in-depth processing capabilities Traditional solutions are mostly based on simple web crawling and template splicing, lacking multimodal and multi-layered semantic understanding and content reconstruction capabilities. The retrieved information is mostly public and shallow material, without in-depth semantic analysis and business context reconstruction, resulting in highly repetitive report content, lack of industry-specific analysis, and difficulty in providing differentiated support in highly competitive decision-making scenarios.
[0011] 2. Insufficient personalization and customization The report lacks adaptive generation capabilities to different industry knowledge structures, organizational requirements, and expression styles, and cannot accurately match and dynamically optimize professional terminology, industry standards, and report chapter structures. When combined with the company's private knowledge base, it lacks targeted extraction and weighting strategies, resulting in a low degree of fit between the results and the actual business context, which seriously affects the usability of the report.
[0012] 3. Limited coverage and accuracy of information recall. Most existing solutions rely on a single knowledge source (such as an internal corporate database or a public web search), which limits the scope of retrieval. Furthermore, the retrieval methods are based on keywords or shallow matching, which cannot meet the needs of capturing comprehensive information under complex business semantic relationships. The lack of a dynamic processing mechanism for timeliness factors makes it difficult to reflect hot information, the latest market changes, and emergencies in the reports in a timely manner.
[0013] 4. Loose generation logic and lack of consistency control There is a lack of a unified quality control mechanism for processing recall results, such as cross-chapter logic verification, contextual consistency review, and data reference verification. The large model generation results are at risk of "hallucination," which may produce conclusions that do not conform to the facts or conflicting viewpoints. Due to the lack of a rigorous post-generation verification mechanism, this problem will directly impact the accuracy and professionalism of the report.
[0014] 5. Insufficient scalability and cross-environment adaptability The current system requires extensive manual adaptation when introducing large language models of different types or scales. The interfaces and modules are highly coupled, the deployment environment is limited, and it is difficult to smoothly migrate to the domestic information technology innovation ecosystem or hybrid cloud architecture. It does not have the high concurrency and multi-task parallel processing capabilities for TB-level corpora, which limits its application potential in large enterprise groups or high-frequency report generation scenarios.
[0015] 6. Weak quality traceability makes effective auditing and tracing difficult. Traditional solutions lack end-to-end log collection and data traceability mechanisms. Once the generated content is incorrect or controversial, it is impossible to quickly locate the source of information, recall strategies, and model reasoning process. This problem is particularly prominent in highly sensitive fields such as finance, compliance, and healthcare, and may directly affect the company's compliant operations and decision-making security.
[0016] Therefore, the present invention sets the following technical objectives: 1. Construct a research report rapid generation technology architecture that combines artificial intelligence and multi-path vector recall, which can efficiently recall, deeply fuse and enhance the generation of multi-source heterogeneous data; 2. While ensuring originality, professionalism, and logical consistency, significantly improve the customization and business fit of the report; 3. It has high scalability, cross-environment adaptability and full-chain traceability, and can be widely used in a variety of large-scale business scenarios such as market research, strategic decision-making and industry analysis; 4. Through component-based and containerized deployment, it supports domestic adaptation and large-scale parallel processing, reducing operation and maintenance costs and improving system flexibility.
[0017] The technical architecture of this invention consists of the following modules.
[0018] 1. Task Configuration and Interaction Module It supports inputting the report topic, the target large language model type to be called (such as Qwen-7B, Baichuan-13B, Yi-34B, DeepSeek-R1, Qwen3-32B, etc.), the estimated word count, and industry templates (the system's built-in report structure reference framework widely used by China International Engineering Consulting Corporation); it provides parameterized business constraints and formatted output rules configuration to meet the needs of different projects.
[0019] 2. Data Processing and Knowledge Storage Module This module is responsible for data preprocessing, knowledge vectorization storage, and semantic retrieval index construction for long documents. It is the core foundation for knowledge-driven generation in the system. Its main function is to build a searchable and traceable private knowledge index, and it does not directly participate in the report generation process.
[0020] (1) Definition of the source and function of long documents The long documents processed by the system mainly originate from the company's internal knowledge assets, such as feasibility study reports, project evaluation documents, and industry analysis briefs from previous years. After structured fragmentation and semantic vectorization, these documents form the company's private knowledge base, providing semantic indexing support for subsequent multi-source retrieval modules.
[0021] The output of this module only includes standardized semantic vectors and their metadata indexes, which serve as the foundation of the system's knowledge layer and do not participate in the actual generation logic.
[0022] (2) Hybrid semantic segmentation and vectorization process The system adopts a hybrid semantic segmentation strategy, which combines semantic boundary detection and structured tag parsing algorithms to achieve document-level semantic segmentation and high-dimensional vector mapping.
[0023] Semantic Boundary Detection By using a semantic flow density model based on the Transformer architecture to identify text breakpoints and thematic transitions, continuous segmentation of natural paragraphs can be achieved.
[0024] Structured Tag Parsing Parse the chapter titles, serial numbers, tables, and metadata tags in the original report to establish the logical hierarchy and relationship identifiers of the content.
[0025] The two-stage results are normalized by the Semantic Fusion algorithm to generate a standardized set of fragments with continuous semantics and structural hierarchy.
[0026] (3) Vectorization mapping and data entry mechanism Each semantic fragment is mapped to a high-dimensional dense semantic vector through a text embedding model (such as bge-m3 or text-embedding-ada-002).
[0027] The system attaches metadata information (including document identifier, chapter index, industry tag, timestamp, source path, etc.) to each vector and stores all vectors in a vector database (such as Milvus or Qdrant) to form a highly scalable index structure.
[0028] This vector database forms the underlying layer of the subsequent "multi-path vector recall module," providing a semantic query and matching foundation for the recall phase.
[0029] (4) Knowledge base maintenance and dynamic updates When new industry documents or research materials enter the system, knowledge base expansion can be achieved simply by performing incremental vectorization processing in this module. This module maintains database consistency and version logs, and does not participate in the dynamic write-back and report content updates during the generation process. Subsequent knowledge reverse updates are uniformly scheduled by the "multi-path recall module." After semantic vectorization and database entry are completed, the system can enter the multi-source retrieval stage. The following modules, with the vector database as their core, perform cross-source, multi-path semantic matching to support the generation enhancement process.
[0030] 3. Multi-path vector recall module This module is responsible for generating a query vector based on user input and system configuration during the report generation task. Using this vector as the core retrieval basis, it retrieves semantic information highly relevant to the report topic from the aforementioned vector database and other knowledge sources in parallel. This module achieves cross-source fusion and multi-channel semantic matching between private knowledge bases, public knowledge sources, and real-time network resources, and is a key component in realizing the Retrieval-Augmented Generation (RAG) mechanism of this invention.
[0031] (1) Recall logic and query vector generation mechanism Input source and generation basis The recall process in this module is driven by task parameters provided by Module 1 (Task Configuration and Interaction Module). Input information includes report topic, industry domain tags, target large language model type, keyword constraints, and chapter structure task identifier.
[0032] Based on these parameters, the system uses a text embedding model to transform the report topic and context description into one or more query semantic vectors.
[0033] These query vectors represent the core topic distribution of the user task in the semantic space and are used to match them with the stored knowledge vectors.
[0034] The role of retrieval target and vector database The vector database built in Module 2 stores long document fragment vectors such as enterprise internal feasibility study reports and project analysis documents that have been semantically segmented and vectorized. These constitute the system's private industry knowledge base channel.
[0035] During the multi-path recall phase, Module 3 first uses the generated query vector as the key to perform a similarity search on the vector database, and selects the set of internal knowledge fragment vectors that are most relevant to the report topic by calculating cosine similarity or vector distance.
[0036] (2) Cross-source multi-path retrieval and semantic fusion process To ensure the coverage and timeliness of the recalled content, this module performs three parallel searches simultaneously: Private industry knowledge base retrieval channel: Retrieve the semantic vector set corresponding to the annual reports of China Consulting Corporation from the vector database established in Module 2, and recall the content fragment with the highest matching degree with the current task topic.
[0037] Public general knowledge base retrieval channel: Calls the system's pre-built open knowledge sources (such as encyclopedia datasets and industry white paper indexes), performs the same query vector matching process, and expands the background information and general knowledge support required for report generation.
[0038] Real-time online resource retrieval channel: By calling external APIs or search engine interfaces, a query request is generated based on the report topic and keywords to obtain the latest industry news, policy information and dynamic data from real-time online resources.
[0039] The obtained text is quickly semantically embedded to form a temporary vector cache, which is then used in the fusion stage along with the other two recall results.
[0040] (3) Multi-path fusion and sorting optimization The three-way recall results undergo semantic fusion and re-ranking in this module, with the following processing logic: A unified vector pool is established for the recall results. The ranking is based on a weighted fusion algorithm, considering similarity scores, timeliness weight, knowledge source credibility, and content diversity. The Reranker model is applied to perform multi-dimensional re-ranking of candidate results to ensure that the content input to the generation module takes into account both relevance and information freshness. Deduplication filtering removes semantically repetitive or textually redundant segments.
[0041] The merged highly relevant segments will be returned to Module 4 (the generation enhancement module) in text form for chapter-level Prompt building and report generation.
[0042] (4) Mechanism for connecting and reusing vector databases In this stage, the vector database not only serves as the knowledge foundation for the retrieval index, but also plays a role in providing semantic context guidance. When complementary fragments on the same topic are recalled, the system can identify potential relationships through vector proximity analysis to guide the argumentation structure of the generated report.
[0043] The vector database supports reverse update of recall results, which means that new semantic embeddings can be generated in the generated report text and written back to the database to continuously optimize the knowledge vector space and realize knowledge evolution.
[0044] This means that Module 2 and Module 3 form a closed loop: Module 2 is responsible for knowledge accumulation and index construction, Module 3 is responsible for semantic retrieval and dynamic application of knowledge, and the subsequently generated content can incrementally update the knowledge base of Module 2, realizing the system's self-learning and iterative enhancement.
[0045] 4. Generate enhancement modules The prompt is built chapter by chapter, combining the recalled content with industry templates, and a large language model is used to generate an initial draft. Global optimizations are performed to ensure syntactic consistency, logical coherence, and business fit, eliminating fundamental errors.
[0046] 5. Document Output and Log Management Module Supports exporting to multiple formats including PDF, WPS, and DOCX, and provides online preview. Completely records generation task parameters, runtime, data source, and model call logs, enabling traceability and quality auditing.
[0047] The technical solution of this invention is as follows: A report generation method combining artificial intelligence and multi-path vector recall includes the following steps: 1) The task configuration and exchange module receives the report topic, selected target large language model, selected industry template and generation parameters input by the user; it transforms the constraints in the generation parameters into control parameters and embeds them into the selected industry template to form constraint vectorization features, which are used to guide the generation process to maintain business matching in terms of text logic, terminology consistency and expression depth. 2) The selected target large language model is invoked to generate a chapter-level structured report outline based on the constrained vectorized features and the input report topic; 3) Process each chapter in the chapter-level structured report outline using steps a) to e): a) Generate chapter query vectors based on chapter titles and keywords, and retrieve internal knowledge fragments related to the chapter from the enterprise's vector database based on the chapter query vectors, and store them as recall vectors in a temporary database; b) Based on the chapter query vector of the chapter, retrieve information related to the title of the chapter from the public knowledge base, embed it into vectorized processing, generate a recall vector, and store it in a temporary database; c) Based on the keywords of this chapter, obtain industry data related to the topic report from public online resources, embed and vectorize it, generate recall vectors, and store them in a temporary database; d) Fusion of recall vectors in the temporary database: First, for any two recall vectors in the temporary database, if their similarity is higher than a set threshold and their topic tags are the same, mark one of the recall vectors as a redundant item and remove it; then calculate the semantic similarity between the chapter query vector of the chapter and each of the aforementioned recall vectors, and delete recall vectors whose semantic similarity is lower than the set similarity threshold or whose semantics deviate from the title of the chapter; then determine the comprehensive relevance between the corresponding recall vector and the chapter based on the semantic similarity, source credibility, and topic coverage of each retained recall vector and the chapter query vector of the chapter, and output the optimal recall result set of the chapter based on the comprehensive relevance; e) Based on the optimal recall result set of the chapter and combined with the title and business constraints of the chapter, construct the chapter-level generation prompt for the chapter and input it into the target large language model to generate the chapter text and analysis content. 4) Based on the chapter-level structured report outline, combine the chapter text and analysis content of each chapter obtained in step 3) to form the initial draft of the report; then conduct logical review, data citation verification and terminology consistency correction on the initial draft of the report to obtain the final report.
[0048] Preferably, the generation parameters include the number of chapters in the report and business constraints.
[0049] Preferably, the business constraints include language style, report structure, reference time period, industry terminology standards, and compliance restrictions.
[0050] Preferably, in step 3), the processing of each chapter is executed in parallel in a distributed environment.
[0051] Preferably, the method for constructing the vector database is as follows: semantic segmentation and vectorization are performed on the selected enterprise's internal feasibility study reports and project analysis documents to obtain multiple document fragment vectors; then, metadata information is added to each document fragment vector to generate internal knowledge fragments; the metadata information includes document identifier, chapter index, industry tag, timestamp, and source path.
[0052] Preferably, the chapter-level generation prompt includes chapter titles and topics, keywords and argument summaries, as well as auxiliary knowledge fragments extracted based on the recall results and dynamic business constraint embedding vectors.
[0053] Preferably, each chapter in the chapter-level structured report outline includes a title, core arguments, several keywords, and topic weights.
[0054] Compared with existing technologies, this invention has significant technical advantages and application value in many aspects.
[0055] 1. The comprehensiveness and accuracy of information retrieval have been significantly improved. By parallel recall and semantic fusion of three types of data sources—private industry knowledge bases, public general knowledge bases, and real-time online retrieval—this invention can significantly improve knowledge coverage and information freshness in a single task, avoiding the content omissions and timeliness issues that are prone to occur in traditional single-source retrieval. The invention also introduces a multi-dimensional re-ranking mechanism based on the Reranker model, which not only matches based on similarity but also incorporates a timeliness weight and coverage balancing factor to optimize the recall results in terms of accuracy, comprehensiveness, and timeliness.
[0056] 2. The structure and logical rigor of the generated results are significantly enhanced. By leveraging a chapter-level Prompt design and generation enhancement mechanism, the Search Enhancement Generation (RAG) strategy is refined to the chapter level, ensuring consistency and professionalism in content coverage, language expression, and logical deduction for each chapter. Furthermore, the integration of full-text consistency verification ensures that the generated results meet high-quality standards in terms of argumentation chains, data citations, and industry terminology uniformity, avoiding issues such as content fragmentation, information conflicts, or conceptual misuse.
[0057] 3. High degree of customization and cross-industry adaptability Based on an industry template-driven generation method, it can be quickly adapted to multiple vertical fields such as finance, healthcare, manufacturing, energy, and the Internet, and supports the customization of private knowledge bases for specific enterprises, achieving deep integration of business logic and professional vocabulary system; the templates and knowledge bases are independently decoupled, which facilitates cross-scenario migration and quick switching of industry backgrounds, reducing secondary development and deployment costs.
[0058] 4. Excellent performance and scalability, supporting large-scale parallel tasks. Leveraging a modular architecture and containerized deployment, the system can run stably in private cloud, hybrid cloud, and domestic IT innovation environments, and supports horizontal elastic scaling to meet the needs of rapid generation of batch reports. It supports seamless switching between various domestic large language models and open-source models (such as Qwen, Baichuan, Yi series, etc.), which can meet the needs of different parameter scales and inference speeds, and can flexibly choose between performance and cost according to business scenarios.
[0059] 5. Traceable and auditable end-to-end quality management By logging the entire process and recording the retrieval process, the entire process of generating the task is traceable. From the original retrieved data to the final draft, the source and processing logic can be tracked, ensuring the credibility and compliance of the content. In high-reliability scenarios such as industry compliance audits, scientific research result verification, and business negotiations, this advantage can significantly improve the usability and credibility of the report.
[0060] 6. Significantly reduces labor costs and shortens production cycles. The system can generate well-structured and highly relevant industry research reports within minutes, significantly shortening the time required for traditional manual research, data compilation, and report writing—a process that typically takes days or even weeks. While ensuring content quality and logical depth, it effectively reduces human resource investment and optimizes the efficiency of enterprise knowledge production and decision support.
[0061] 7. Enhance the value and utilization rate of enterprise knowledge assets The integration of private knowledge bases and public knowledge graphs enables the structured utilization and dynamic replenishment of enterprise-accumulated data, avoiding the phenomenon of knowledge silos. Through semantic indexing and enhanced retrieval, unstructured texts accumulated within the enterprise (such as historical reports, meeting minutes, and business files) can be fully mined and reorganized to form a continuously evolving knowledge-driven system.
[0062] In summary, this invention not only overcomes the problems of insufficient recall and logical inconsistencies in research report generation at the technical implementation level, but also achieves a balance between information coverage and business customization in terms of application value. It has broad industry application prospects and significant commercial value, and can play an important role in highly competitive and decision-sensitive business environments. Attached Figure Description
[0063] Figure 1 This is a schematic diagram of the overall system architecture of the present invention.
[0064] Figure 2 This is a flowchart of the multi-vector recall fusion process.
[0065] Figure 3 Generate an overall flowchart for the report. Detailed Implementation
[0066] The present invention will now be described in further detail with reference to the accompanying drawings. The examples given are only for explaining the present invention and are not intended to limit the scope of the present invention.
[0067] (I) System Overall Architecture The research report rapid generation system of this invention adopts a B / S (Browser / Server) distributed architecture design, which consists of three core subsystems: front-end task configuration and interaction layer, back-end knowledge retrieval and generation reasoning layer, and storage and resource management layer. At the same time, it achieves low coupling and high scalability between modules through microservice design (as shown in Figure 1).
[0068] 1. Front-end task configuration and interaction layer Task configuration interface: Implemented based on a responsive web front-end framework (such as Vue.js or React), compatible with multi-terminal access. It provides options for report topic input, industry template selection, target large language model (LLM) selection, and generation parameter settings (word count, number of chapters, business constraints, etc.).
[0069] Interaction logic layer: Supports real-time validation of task parameters (type validation, business rule validation), and sends task requests to the backend task management module via WebSocket or HTTP API.
[0070] Visual preview and export: The report generation process displays a real-time progress bar, preview of review recall information, and chapter status. The final generated report can be rendered online and exported to standard formats such as PDF, WPS, and DOCX.
[0071] 2. Backend Knowledge Retrieval and Generation Reasoning Layer The backend adopts a microservice architecture, which mainly consists of the following logical sub-modules.
[0072] 2.1 Task Scheduling and Control Module It is responsible for receiving configuration parameters submitted by the front end and dynamically scheduling tasks based on task priority, model inference load, and resource availability. In high-concurrency scenarios, it ensures task processing order and system stability through task queues and distributed message middleware (such as RabbitMQ and Kafka).
[0073] 2.2 Knowledge Retrieval and Multi-path Vector Recall Module This module undertakes the core functions of backend knowledge retrieval and semantic recall. It achieves the fusion of cross-source data and the acquisition of highly relevant information through a three-channel semantic retrieval subsystem that runs in parallel. The retrieval process of the module is driven by the generation task parameters provided by the task configuration layer and relies on the aforementioned semantic vector database and real-time data source.
[0074] (1) Input logic and driving information The retrieval process of the three-channel retrieval submodule all takes the structured task information provided by the front-end task configuration and interaction module (module 1) as input, including: Report Topic and Key Terms; Industry Tag; Report chapter hierarchy information (Chapter Index / Section Structure); Temporal Constraints and Date Ranges. Model call parameters and generation constraints.
[0075] Based on the task information mentioned above, the system invokes the embedding model to generate one or more query semantic vectors. This vector represents the semantic feature distribution of the current report topic and serves as the core input for the three sub-modules to perform retrieval operations.
[0076] (2) Logic of the three-channel retrieval submodule Private Knowledge Retrieval Submodule Input: Report topic, industry tags, and query semantic vector; Data Source: The vector database constructed by Module 2 is formed by the system through hybrid semantic segmentation of long internal documents of China International Engineering Consulting Corporation (CIECC), such as feasibility study reports, project analysis documents, and internal assessment data from previous years. Specifically, Module 2 first performs semantic boundary detection and structured tag parsing on the documents, splitting the original long documents into multiple document fragments with complete semantic and structural hierarchies. Then, it calls a text embedding model to generate corresponding high-dimensional semantic vectors for each document fragment, and adds metadata information such as document identifier, chapter index, industry tag, timestamp, and source path to each vector. Finally, the semantic vectors and their metadata are uniformly stored in the vector database, forming a searchable and traceable set of enterprise private knowledge vectors.
[0077] Retrieval logic: Retrieve the semantic fragments that best match the query vector from the database by calculating cosine similarity or vector distance, and obtain highly relevant internal knowledge content; Output: A set of highly relevant private knowledge fragments, including relevant metadata (source document ID, topic tags, timestamps, etc.).
[0078] Public Knowledge Retrieval Submodule Input: Report topic, keywords, and query semantic vector; Data source: Pre-built open knowledge bases (such as encyclopedia datasets, industry white paper indexes, academic paper abstracts, etc.); Retrieval logic: Perform semantic matching with the same query vector to recall more general knowledge content with broader coverage, in order to supplement the deficiencies of industry-specific private knowledge bases; Output: A collection of Generic Knowledge Fragments, used for background information, industry comparisons, and cross-domain references in the report.
[0079] Real-Time Web Retrieval Submodule Input: Report topic keywords and timeliness constraints; Data source: Real-time API interface or search engine data source, used to obtain the latest policy updates, news information and market data; Retrieval logic: Generate a query request (Textual Query) based on the keywords in the task configuration, obtain the latest text, perform fast embedding calculation and generate a real-time semantic vector, and match it with the query vector to filter out highly relevant results; Output: A collection of real-time semantic fragments, primarily used to enhance the timeliness and data update dimensions of reports.
[0080] (3) Integration and Result Output Mechanism The recall results output by the three-channel submodules enter the unified fusion processing unit and perform the following steps: Semantic Vector Fusion: Merges relevant vector sets from different data sources into a unified vector pool; Weighted Re-Ranking: Optimizes ranking based on similarity score, timeliness weight, and source credibility; Deduplication and Consistency Filtering: Removes duplicate content and checks the consistency of information across sources; Output: The final generated highly relevant text and its metadata are passed to the generation enhancement module for chapter-level Prompt construction and report generation tasks.
[0081] (4) Explanation of inter-module relationships and data connection A clear data transfer chain is formed between the three-channel retrieval module and the front-end task configuration layer and data processing layer: Module 1 is responsible for generating task parameters and topic vectors; Module 2 is responsible for knowledge data import and semantic vector storage; Module 3 performs retrieval and fusion from Module 2's private database and other external sources based on the query vector from Module 1.
[0082] This design ensures consistency across the entire system's recall logic, from task definition to semantic matching, and maintains a high degree of coupling between search results and report topics, structure, and timeliness.
[0083] 2.3 Hybrid Semantic Segmentation and Vectorization Module This module is a crucial step in the system's knowledge standardization and semantic index construction. It is responsible for semantic segmentation, vectorization mapping, and database processing of the retrieved long documents or raw data content to support efficient semantic retrieval and generation tasks in subsequent stages.
[0084] (1) Input object and long document definition The input processed by this module mainly comes from the recall result set of the knowledge retrieval and multi-path vector recall module 2.2.
[0085] In this module, after retrieval from a private industry knowledge base, a public general knowledge base, and real-time online resources, the system obtains multiple original text or document objects related to the report topic. These documents are typically long, well-structured content, belonging to the long document type, for example: Industry feasibility study report, annual market analysis report, expert evaluation report; Open source white papers, statistical bulletins, and original policy documents; News summaries or data description documents obtained through online searches.
[0086] Therefore, the processing object of this module is indeed the "long document set" retrieved in the previous stage. The system performs hybrid semantic segmentation and vectorization processing on each long document in turn.
[0087] (2) Long document decomposition and semantically related fragment generation For each long input document, this module employs a hybrid semantic segmentation strategy, combining semantic boundary detection and structured tag parsing algorithms to decompose the long document into a series of semantically associated text fragments. Semantic Boundary Detection: Based on the semantic flow density model, the system automatically identifies semantic breakpoints and topic transitions in the document, ensuring that the fragmentation process does not disrupt semantic coherence.
[0088] Structured Tag Parsing: The document's internal structured markers (such as heading levels, table labels, time markers, and summary descriptions) are parsed and redefined to extract logical structure boundaries.
[0089] Fusion Normalization: The two parsing results are aggregated based on semantic consistency to generate a set of semantic fragments with clear topic affiliation and contextual coherence. Each fragment contains a traceable document source identifier (Document ID) and a section index (Section Index).
[0090] (3) Vectorization mapping and storage method The system calls the embedding model (e.g., bge-m3, text-embedding-ada-002) for each of the above semantic segments to generate the corresponding high-dimensional dense semantic vector.
[0091] Each vector group is associated with its corresponding segment, and metadata (source_doc_id, section_index, timestamp, industry_tag, relevance_score) is recorded. The system then writes the generated semantic vectors and metadata into a vector database (such as Milvus or Qdrant).
[0092] All semantic fragment vectors within the same long document are stored in a document-level vector index structure for subsequent hierarchical semantic retrieval, contextual tracing, and chapter matching.
[0093] (4) Vector retrieval and application logic in subsequent stages Once the vector database is built, its contents will be directly accessed in subsequent stages of the system.
[0094] Module 3 (Multi-path Vector Recall Module) Phase When the report generation task begins, the system generates a query semantic vector by embedding the report topic, keywords and industry tag information input by module 1 (task configuration layer).
[0095] Module 3 retrieves matching high-dimensional dense semantic vectors from the vector database generated and stored in this module based on the query vector.
[0096] Matching is typically calculated using cosine similarity or vector distance to obtain the most relevant set of segments. These segments are then fused and sorted before being input into the report generation module (Module 4) to construct the chapter-level Prompt.
[0097] Module 4 (Generation Enhancement and Consistency Verification Module) Phase When generating the report, the system automatically calls the corresponding set of semantic fragment vectors in the vector database according to the chapter topic, and performs context augmentation and argument enhancement operations to ensure that the generated content is consistent with the existing knowledge semantics.
[0098] Therefore, the semantic vector data generated by this module plays a dual role in the overall system process: Provides a high-precision indexing foundation for semantic retrieval (used in Module 3); Provides semantic support and consistency reference for the generation of enhancements (used by Module 4).
[0099] (5) Explanation of module connection and data flow relationship From a system perspective, the relationships between the modules are as follows: Module 2.2 outputs the long document retrieved; Module 2.3 performs fragmentation, embedding, and storage of these long documents; The vector database serves as the knowledge base for the retrieval operations in Module 3; Module 3 generates a query vector based on the task parameters of Module 1 and performs a retrieval in the vector database; The retrieved vector fragments are used as input for the generation and enhancement stage.
[0100] This hierarchical design enables the system to form a clear data processing chain: The process of "retrieving long documents → segmenting into fragments → vectorizing and storing in the database → querying and retrieving → content generation" ensures logical continuity and traceability in each step.
[0101] 2.4 Generate Enhancement and Consistency Verification Module The system performs chapter-level Prompt construction on the recalled content and calls the selected large language model (such as Qwen, Baichuan, Yi series, etc.) to generate the initial draft text.
[0102] After generation, a post-generation consistency check mechanism is used to detect data conflicts, logical breaks, and factual errors, and secondary corrections are made in conjunction with the business template.
[0103] 2.5 Log and Monitoring Module A distributed logging system (such as ELK Stack) is used to monitor and record the entire generation process. The log content covers data source access records, retrieval time, model call count, and inference result summary.
[0104] 3. Storage and Resource Management Layer This layer is designed to securely and efficiently manage large-scale corpora, vector indexes, metadata, and intermediate generated results.
[0105] Relational Database Management System (RDBMS): Based on MySQL or PostgreSQL, it stores task parameters, industry templates, configuration files, and system metadata.
[0106] Vector DB: It uses Milvus, Qdrant or FAISS to build high-dimensional semantic indexes, supports distributed and sharded deployment, and enables fast retrieval of hundreds of millions of documents.
[0107] Object storage system: Based on MinIO or Ceph, it provides scalable storage for intermediate files, exported results, and long text source data during the report generation process.
[0108] Caching layer: Use Redis or Memcached to cache frequently accessed search results and template fragments to reduce response latency.
[0109] 4. System security and adaptability design Data security: The entire transmission chain uses the TLS encrypted transmission protocol, and access to the private knowledge base requires verification through a fine-grained access control mechanism based on OAuth 2.0.
[0110] Domestic production and cross-platform support: It supports operation on domestic CPUs (such as Phytium and Kunpeng) and domestic GPUs (such as Ascend and Cambricon), is compatible with mainstream Linux distributions, and can be smoothly migrated to domestic cloud platforms.
[0111] Elastic scaling: Kubernetes enables container orchestration and automatic resource scaling, dynamically allocating computing instances based on task queue length and load conditions to achieve high-concurrency task processing capabilities.
[0112] (ii) System workflow (e.g.) Figure 3 (As shown) The report generation system of this invention adopts a hierarchical driving and parallel processing mode, and the overall workflow is as follows: 1. Task initialization and parameter configuration Users enter the report topic in the interactive interface, select the target Large Language Model (LLM) version and industry template, and set the word count range and business constraints.
[0113] Business constraints fall under the category of parameters that can be freely set in this field, such as language style, report structure, reference time range, industry terminology standards, and compliance restrictions.
[0114] The innovation of this invention lies in the dynamic semantic mapping mechanism of constraints. The system first performs structured parsing of business constraints, classifying them into different constraint types such as chapter structure, terminology specifications, industry style, and content depth. Then, based on the constraint type, it constructs corresponding constraint semantic description templates and calls a text embedding model to map these constraint descriptions into constraint vectors within the same semantic space as the generated text. The system further weights and fuses multiple constraint vectors according to chapter attributes and business priorities, generating unified constraint vectorized features as control parameters during the generation process. In the chapter-level Prompt construction stage, these constraint vectorized features are embedded into the generation guidance instructions through positional injection, thereby continuously constraining chapter logic, terminology selection, and expression depth throughout the generation process, achieving a high degree of consistency between the generated content and business requirements.
[0115] Compared with the traditional parameter setting method, this semantic layer constraint embedding method has the following technical advantages: This ensures that the generated reports adhere to business logic and style requirements in both content and expression; It supports the parallel application of multi-dimensional constraints, enabling coordinated control of format, tone, and professionalism; Improve the contextual consistency of the generated results and the logical coherence between chapters.
[0116] 2. Report Outline Generation The system calls the selected large language model and generates a chapter-level structured outline based on the industry template (a built-in report structure reference framework widely used by China Consulting Corporation) and the input report topic.
[0117] Each chapter includes a title, core arguments, several keywords, and topic weights. These keywords are general information extraction results, which are common technologies and are automatically generated through semantic analysis or Keyphrase Extraction algorithms.
[0118] In the next phase, the system will use these chapter keywords and topic vectors as the basis for recall queries.
[0119] 3. Chapter-level parallel processing The processing of each chapter is executed in parallel in a distributed environment, and the specific steps are as follows: a. Semantic Recall of Private Industry Knowledge Base Based on the chapter title, keywords, and their semantic vector (a chapter query vector generated by jointly embedding the chapter title and keywords), a retrieval operation is performed from the vector database built in Module 2 to retrieve internal knowledge fragments that are highly relevant to the chapter.
[0120] This operation enables semantic matching between chapters and enterprise proprietary knowledge (such as feasibility study reports from China International Engineering Consulting Corporation over the years).
[0121] b. Semantic Recall of Public Knowledge Base Based on the same chapter query vector, relevant materials (industry white papers, statistical data, policy documents, etc.) related to the chapter titles are retrieved from public knowledge bases to supplement general information and improve content coverage.
[0122] c. Real-time online resource recall Based on chapter keywords and semantic vectors, the system invokes online search interfaces or web crawling tools to retrieve industry trends and data updates related to the current report chapter's theme from publicly available online resources. To ensure the timeliness of the retrieved content, the system first filters the search results based on a preset time window and analyzes the content's publication or most recent update time. Then, it calculates a timeliness score for candidate content that meets the time constraints and weights this score with the chapter's semantic similarity score, retaining only text content that simultaneously meets both semantic relevance and timeliness thresholds to determine the latest industry information used for report generation.
[0123] The search results are quickly embedded and vectorized before being temporarily stored in a temporary database.
[0124] d. Multi-channel recall fusion and reordering When the three recall results enter the fusion stage, the system executes the following algorithm steps: Redundancy Detection: Using textual similarity and vector proximity analysis, each segment is compared item by item. If the cosine similarity is higher than a set threshold (usually 0.92) and is consistent with the chapter title label, it is marked as a redundant item and removed.
[0125] Low-Relevance Filtering: Sort by semantic similarity score between chapter query vector and recall vector (timeliness weight can be incorporated). If the score is below the threshold (e.g., <0.6) or the semantics deviate from the chapter title, it is removed.
[0126] In the relevance calculation and weighted ranking stage, the system constructs a multi-dimensional feature score for each recalled semantic vector. Specifically, it first calculates the semantic similarity between the recalled vector and the corresponding chapter query vector; secondly, it calculates the timeliness score based on the publication or update time of the recalled content; simultaneously, it determines the source credibility score based on the data source type of the recalled content and its pre-set credibility level table; and it analyzes the coverage of the recalled content to the core theme of the chapter by combining chapter keywords and sub-topic vectors to obtain a theme coverage score. After normalizing the above multi-dimensional scores, the system calculates the comprehensive relevance score between the recalled vector and the chapter theme using a weighted fusion algorithm, and performs re-ranking and filtering accordingly, finally outputting the recall result set with the best comprehensive relevance.
[0127] e. Build chapter-level Prompts and generate report text. The system uses the fused recall results as chapter knowledge input, and constructs a chapter-level Prompt (Chinese name: chapter-level generation guidance instruction) by combining the chapter topic and business constraints. This Prompt includes: Chapter titles and themes; Keywords and abstract of arguments; Auxiliary knowledge fragments extracted from the recall results; Dynamic business constraint embedding vector.
[0128] After calling the large language model, the model performs retrieval enhancement generation (RAG) based on the prompt for each chapter, automatically generating the chapter text and analysis content, which are then combined to form the initial draft of the report.
[0129] 4. Overall consistency and optimization All generated chapters are processed by the Post-Generation Consistency Optimization module, which performs full-text logical review, data citation verification, and terminology consistency correction.
[0130] Perform global vector consistency comparison at the semantic level to ensure the connection between themes across chapters and the closure of the argument chain.
[0131] This step ensures that the report is fluent, logically consistent, and conforms to business expression standards.
[0132] 5. Document Export and Log Recording The system supports exporting reports in formats such as PDF, WPS, and DOCX, while also generating a full-text visual preview.
[0133] The entire process records task parameters, data source recall, generation time, model call logs, and version tracking to achieve auditing and traceability management.
[0134] (III) Key Technical Details 1. Hybrid Semantic Segmentation Strategy This strategy is an innovative algorithm module of the present invention, used to process long documents and generate high-precision semantic fragment indexes.
[0135] The innovation lies in the cross-layer semantic recognition and structured fusion process, specifically implemented as follows: Perform semantic boundary detection on long documents and identify semantic breakpoints through a deep semantic flow density analysis model; Based on Structured Tag Parsing, key elements such as chapter titles, tables, and time stamps are logically divided into domains. The segmentation results are aggregated according to semantic relevance using a semantic fusion algorithm to generate "semantically associated fragments". The embedding model is invoked to generate high-dimensional dense semantic vectors for each segment, and stored in a vector database to form a document-level index structure.
[0136] This innovative strategy avoids the semantic fragmentation problem of traditional sharding algorithms through a two-layer detection and fusion mechanism, which can significantly improve recall accuracy and context continuity, and improve the system's retrieval and generation quality.
[0137] 2. Multi-path vector recall fusion mechanism The system implements parallel retrieval through three channels: a private industry knowledge base, a public general knowledge base, and real-time networked resources.
[0138] The innovation lies in the introduction of a Dynamic Weighted Fusion Algorithm, which simultaneously considers semantic similarity score, temporal weight, content diversity, and source trust factor in dimensionality calculation.
[0139] This algorithm achieves adaptive balancing of multi-path results in a high-dimensional semantic space, improving the comprehensiveness and effectiveness of recall.
[0140] 3. Generation Enhancement and Consistency Optimization The generation enhancement module of this invention is not a typical recall process, but rather employs a two-stage consistency enhancement mechanism: Section-level Semantic Enhancement: Before generating a section, perform Vector Aggregation Analysis on the recall results to extract core semantic features and enhance the context.
[0141] Document-level Logical Consistency Optimization: After generation, cross-section vector validation is used to verify the consistency of terms and argument chains, and semantic drift and factual conflicts are automatically corrected.
[0142] This mechanism achieves closed-loop control of "recall-fusion-generation-verification", significantly improving the accuracy and business adaptability of the generated content.
[0143] 4. Localization and cross-platform adaptation The system supports operation on domestic CPUs (such as Kunpeng and Phytium) and GPUs (such as Ascend and Cambricon), and is compatible with domestic large-scale models such as Qwen, Baichuan, and Yi.
[0144] By optimizing the hardware instruction layer and accelerating vector computation, the performance of the entire inference chain is improved.
[0145] 5. Containerized deployment of scalable architecture Docker and Kubernetes are used to containerize and encapsulate each module component, enabling independent deployment and dynamic horizontal scaling of modules.
[0146] It supports online hot updates and migration to multi-cloud environments, meeting the business needs of enterprises for large-scale parallel report generation.
[0147] Although specific embodiments of the invention have been disclosed for illustrative purposes to aid in understanding and implementing the invention, those skilled in the art will understand that various substitutions, variations, and modifications are possible without departing from the spirit and scope of the invention and the appended claims. Therefore, the invention should not be limited to the content disclosed in the preferred embodiments, and the scope of protection claimed by the invention is defined by the claims.
Claims
1. A report generation method combining artificial intelligence and multi-path vector recall, comprising the following steps: 1) The task configuration and exchange module receives the report topic, selected target large language model, selected industry template and generation parameters input by the user; The constraints in the generated parameters are transformed into control parameters and embedded into the selected industry template to form constraint vectorization features, which are used to guide the generation process to maintain business matching in terms of chapter logic, terminology consistency and expression depth. 2) The selected target large language model is invoked to generate a chapter-level structured report outline based on the constrained vectorized features and the input report topic; 3) Process each chapter in the chapter-level structured report outline using steps a) to e): a) Generate chapter query vectors based on chapter titles and keywords, and retrieve internal knowledge fragments related to the chapter from the enterprise's vector database based on the chapter query vectors, and store them as recall vectors in a temporary database; b) Based on the chapter query vector of the chapter, retrieve information related to the title of the chapter from the public knowledge base, embed it into vectorized processing, generate a recall vector, and store it in a temporary database; c) Based on the keywords of this chapter, obtain industry data related to the topic report from public online resources, embed and vectorize it, generate recall vectors, and store them in a temporary database; d) Fusion of recall vectors in the temporary database: First, for any two recall vectors in the temporary database, if their similarity is higher than a set threshold and their topic tags are the same, mark one of the recall vectors as a redundant item and remove it; then calculate the semantic similarity between the chapter query vector of the chapter and each of the aforementioned recall vectors, and delete recall vectors whose semantic similarity is lower than the set similarity threshold or whose semantics deviate from the title of the chapter; then determine the comprehensive relevance between the corresponding recall vector and the chapter based on the semantic similarity, source credibility, and topic coverage of each retained recall vector and the chapter query vector of the chapter, and output the optimal recall result set of the chapter based on the comprehensive relevance; e) Based on the optimal recall result set of the chapter and combined with the title and business constraints of the chapter, construct the chapter-level generation prompt for the chapter and input it into the target large language model to generate the chapter text and analysis content. 4) Based on the chapter-level structured report outline, combine the chapter text and analysis content of each chapter obtained in step 3) to form the initial draft of the report; then conduct logical review, data citation verification and terminology consistency correction on the initial draft of the report to obtain the final report.
2. The method according to claim 1, characterized in that, The generated parameters include the number of chapters in the report and business constraints.
3. The method according to claim 2, characterized in that, The business constraints include language style, report structure, reference time frame, industry terminology standards, and compliance restrictions.
4. The method according to claim 1, 2, or 3, characterized in that, In step 3), the processing of each chapter is executed in parallel in a distributed environment.
5. The method according to claim 1, 2, or 3, characterized in that, The method for constructing the vector database is as follows: semantic segmentation and vectorization are performed on the selected enterprise's internal feasibility study reports and project analysis documents to obtain multiple document fragment vectors; then, metadata information is attached to each document fragment vector to generate internal knowledge fragments; the metadata information includes document identifier, chapter index, industry tag, timestamp, and source path.
6. The method according to claim 1, 2, or 3, characterized in that, The chapter-level generation prompt includes chapter titles and topics, keywords and argument summaries, as well as auxiliary knowledge fragments extracted based on recall results and dynamic business constraint embedding vectors.
7. The method according to claim 1, 2, or 3, characterized in that, Each chapter in the chapter-level structured report outline includes a title, core arguments, several keywords, and topic weights.