Rule engine-based engineering construction scheme document automatic generation method and system

CN122840016APending Publication Date: 2026-09-29THE SECOND ENG COMPANY OF CCCC FOURTH HARBOR ENG
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202611350423.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-09-02
Publication Date
2026-09-29

AI Technical Summary

Technical Problem

缺陷一:缺乏领域感知的模板-规则匹配机制

Benefits of technology

编制效率大幅提升:将施工方案编制时间从数天缩短至数分钟,覆盖从工程概况到应急处置的多个章节,生成完整度可达85%以上(剩余15%为需人工确认的项目特定数据)。规范引用准确性有保障,避免了人工核对规范的遗漏。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122840016A_ABST
    Figure CN122840016A_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of computer-aided engineering document automatic generation, in particular to an engineering construction scheme document automatic generation method and system based on a rule engine, the method comprising the following steps: in response to a received project name, project field and template path, loading a preset N template rules, the template rules comprising at least a rule identifier and a rule action type; based on a template index and a project data index, constructing a full-text retrieval engine index; extracting project key information from the project data, and sequentially executing the N template rules to generate chapter contents; and based on the template document and the chapter contents, generating an edited version Word document and a pure version Word document. The method of the present application greatly improves the compilation efficiency, guarantees the accuracy of specification references, avoids the omission of manual specification checking, is traceable in quality, and meets different use scenarios.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of computer-aided engineering document automatic generation technology, specifically to a method and system for automatic generation of engineering construction plan documents based on a rule engine. Background Technology

[0002] The preparation of an engineering construction plan is a core step in engineering construction. It requires the generation of a complete technical document based on standardized templates, industry standards, and the actual conditions of the project. This document should cover multiple chapters, including project overview, construction deployment, construction technology, quality and safety assurance measures, environmental protection and civilized construction, personnel allocation, acceptance requirements, and emergency response. Currently, the industry mainly offers the following types of plans:

[0003] (1) Manual compilation mode: Engineering technicians manually compile chapter by chapter according to the standardized template, relying on personal experience to extract relevant data from the project construction organization design (construction organization), which takes several days to several weeks.

[0004] (2) General template filling system: For example, US20150046791A1 discloses a template system for custom document generation, which generates documents by pre-setting placeholders in the template and having the user input or database query replace them. This type of system is suitable for structured form documents such as invoices and contracts, but cannot handle the generation of a large amount of unstructured technical content in solution documents.

[0005] (3) General Automated Document Generation System: US11087219B1 discloses a workflow-based automatic document generation method that triggers content generation through predefined rules. However, this solution is geared towards the general document domain and is not designed for the special needs of the engineering domain, such as chapter dependencies, specification version comparison, and multi-source knowledge base integration.

[0006] (4) Construction information management based on BIM: CN114418330A discloses an automatic matching method for construction resources based on 4D-BIM model, which focuses on the association between the three-dimensional model and the construction progress, and does not involve the automatic compilation of text plan documents.

[0007] Existing technologies still have some shortcomings, such as: Defect 1: Lack of a domain-aware template-rule matching mechanism. Existing template systems cannot automatically identify the correspondence between different chapters in a construction plan template and the compilation rules, still requiring manual judgment of the content source (knowledge base / project data / LLM generation / fixed sentence structure) segment by segment.

[0008] Defect 2: Lack of an execution engine to handle chapter dependencies in the engineering domain. There are complex dependencies between chapters of the construction plan (such as "construction process flow" depending on "construction deployment", "quality and technical measures" depending on "main construction methods"). The existing system does not have the ability to resolve dependencies and perform topology sorting based on directed acyclic graphs (DAGs).

[0009] Defect 3: Lack of a multi-source version comparison mechanism for engineering specifications. The technical specifications (such as JTS, GB, JGJ, etc.) referenced in the construction plan have version evolution issues, and the version of the specifications referenced in the template may be obsolete. The existing system cannot automatically compare the version status of the specifications from the knowledge base, project construction organization, and templates.

[0010] Defect 4: Lack of an LLM content reorganization strategy tailored to the construction field. The general LLM call lacks differentiated prompts for different chapters of the engineering plan, and cannot automatically construct corresponding contexts and instructions based on rule-based action types (extraction and reorganization / project data mining / knowledge base adaptation and integration / flowchart generation). Summary of the Invention

[0011] To address the shortcomings of the existing technologies, the technical problem this invention aims to solve is: how to construct a complete automatic document generation system for the field of engineering construction schemes, capable of automatically identifying the correspondence between template chapters and compilation rules, scheduling execution according to chapter dependencies, cross-comparing standard versions, and reorganizing project data content through a differentiated LLM strategy, ultimately producing two versions of the scheme document: an edited version with source traceability and a clean, unmarked version. Therefore, a method and system for automatically generating engineering construction scheme documents based on a rule engine is proposed.

[0012] To achieve the above-mentioned objectives, the present invention provides the following technical solution: In a first aspect, the present invention provides a method for automatically generating engineering construction plan documents based on a rule engine, comprising the following steps: Receive user input for the project name, project domain, and template path; In response to the received project name, project domain, and template path, a preset template rule is loaded, wherein the template rule includes at least a rule identifier and a rule action type; An index is created for the template documents specified by the template path, generating a template index; Load and index project data, and generate a project data index; Based on the template index and the project data index, a full-text search engine index is constructed; The key information of the project is extracted from the project data according to the full-text search engine index. The key information of the project includes at least the project name, participating units and construction period information. Based on the directed acyclic graph dependency order among the template rules, the template rules are executed sequentially to generate the content of each chapter based on the key project information; Based on the template index, the content of each chapter is filled into the template document to construct an editable Word document, which includes content source tags and specification comparison results; Remove the comment marks from the edited version of the Word document to generate a clean version of the Word document; Output the edited version of the Word document and the clean version of the Word document.

[0013] Furthermore, the rule action type includes at least one of the following: Knowledge base copy and populate is used to copy chapter content from a knowledge base template and perform parameter replacement and specification comparison. Template variable filling is used to fill variables according to a fixed sentence template. Intelligent extraction and rewriting is used to search for project information through a full-text search engine and extract and reorganize it using a large language model; Item-by-item data extraction is used for keyword searching through a full-text search engine and information extraction using a large language model; Framework fusion filling is used to generate data by fusing knowledge base frameworks and project content using a large language model. Intelligent process generation is used to extract process information from upstream chapters and generate flowchart descriptions using a large language model.

[0014] Furthermore, the construction of the full-text search engine index includes: The template document and the project information are segmented using Chinese binary word segmentation technology, and the full-text search engine index is constructed based on the word segmentation results. Import all document content into an SQLite FTS5 virtual table to enable full-text search based on an inverted index; Based on the in-memory dictionary index, retrieve documents categorized by type; Multi-keyword combination search, searching by keyword one by one and merging to remove duplicates.

[0015] Furthermore, the template rules are executed sequentially according to the directed acyclic graph dependency order among the template rules, including: Analyze the dependencies between the template rules and construct a directed acyclic graph; The execution order of each template rule is determined according to the topological sorting of the directed acyclic graph. Each template rule is executed sequentially according to the execution order described above. Once the current rule is completed, the execution of its downstream dependent rules is triggered.

[0016] Furthermore, key project information is extracted from the project data, including: Use a full-text search engine to retrieve keywords related to the project name, participating units, and construction period from the project data; By using a large language model to extract information from the search results, structured key project information is obtained.

[0017] Furthermore, based on the template index, the content of each chapter is filled into the template document to construct an editable Word document, specifically including: Fill the corresponding positions with the content of each chapter according to the chapter structure of the template document; The original cover format, style definition, header and footer settings of the template document are retained.

[0018] Furthermore, after constructing the editable Word document, the process also includes: The edited version of the Word document retains source tags, which are used to indicate the original source of each chapter's content; The standard comparison results are retained in the edited version of the Word document. These results are used to indicate the conformity between the generated chapter content and the preset standards.

[0019] Furthermore, the step of removing comment markers from the edited version of the Word document also includes: Identify template placeholder marks in the template document; Replace the template placeholder with the corresponding actual information in the project's key information.

[0020] Furthermore, it also includes a three-level specification version comparison method to determine whether the generated Word document meets the requirements of the template rules. The specific steps include: Extract the standard numbers from the template rules, use regular expressions to match the preset set of standard prefixes, and extract all referenced standard standard numbers from the template document; The first level of comparison is knowledge base comparison, which involves performing cardinality matching between the extracted specification number and the specification files stored in the specification knowledge base. The cardinality matching ignores the year and version number differences in the specification number and only compares the main body number part of the specification. If there is any version of the specification file with the same main body number as the current specification in the knowledge base, the latest version information of the specification in the knowledge base and the current template reference version information are recorded, and it is determined that the knowledge base has been matched. The second level of comparison is the construction organization comparison. For specifications that fail to match in the knowledge base after the first level of comparison, the referenced specification number is extracted from the project construction organization design document. The specification that fails to match in the knowledge base is compared with the specification number extracted from the construction organization design document. If there is a specification reference record in the construction organization that has the same main specification number as the current specification, the version information referenced in the construction organization is recorded and it is determined that the construction organization has been matched. The third level of comparison is to mark the specification as questionable. For specifications that are not matched after the first and second levels of comparison, they are marked as questionable because none of the three parties have this specification, and the questionable mark is embedded in the generated document.

[0021] Secondly, this invention also provides an automatic generation system for engineering construction plan documents based on a rule engine, including: The input module is used to receive user input of the project name, project domain, and template path; The data processing module is used to load preset template rules in response to the received project name, project domain and template path. The template rules include at least a rule identifier and a rule action type. An index is created for the template documents specified by the template path, generating a template index; Load and index project data, and generate a project data index; Based on the template index and the project data index, a full-text search engine index is constructed; Extract key project information from the project data. The key project information includes at least the project name, participating units, and construction period. Based on the directed acyclic graph dependency order among the template rules, the template rules are executed sequentially to generate the content of each chapter based on the key project information; The output module is used to fill the content of each chapter into the template document based on the template index, and construct an editable Word document, wherein the editable Word document contains content source tags and specification comparison results; Remove the comment marks from the edited version of the Word document to generate a clean version of the Word document; Output the edited version of the Word document and the clean version of the Word document.

[0022] The beneficial effects of this invention are as follows: Significantly improved compilation efficiency: The time for compiling construction plans has been reduced from several days to several minutes, covering multiple chapters from project overview to emergency response, with a completeness rate of over 85% (the remaining 15% is project-specific data requiring manual verification). The accuracy of standard references is guaranteed, avoiding omissions caused by manual verification of standards.

[0023] Quality traceability: Dual-version output mechanism - the edited version retains the source markup and specification comparison results for each piece of content, while the clean version removes all marks and can be delivered directly, meeting different use cases. Attached Figure Description

[0024] Figure 1 This is a flowchart of the method for automatically generating engineering construction plan documents based on a rule engine in Example 1; Figure 2 This is a flowchart of a specific method for automatically generating engineering construction plan documents based on a rule engine, as shown in Example 1. Figure 3 This is a system diagram for automatically generating engineering construction plan documents based on a rule engine, as shown in Example 2. Detailed Implementation

[0025] The present invention will now be described in further detail with reference to specific embodiments. However, this should not be construed as limiting the scope of the present invention to the following embodiments; all technologies implemented based on the content of the present invention fall within the scope of the present invention.

[0026] The preparation of engineering construction plans is a core work process in the field of engineering construction. It requires the generation of complete technical documents covering chapters such as project overview, construction deployment, construction techniques and methods, quality and safety assurance, environmental and civilized construction, personnel allocation, acceptance requirements, and emergency response, based on standardized templates, industry standards, and actual project conditions. Current mainstream solutions include manual chapter-by-chapter writing, general template filling systems, general automated document generation systems, and BIM-based construction information management. However, existing technical solutions generally suffer from three major shortcomings: First, they lack a domain-aware template-rule matching mechanism, failing to automatically identify the correspondence between each chapter and the compilation rules; second, they lack a Directed Acyclic Graph (DAG) execution engine to handle complex dependencies between engineering chapters (such as "construction technology" depending on "construction deployment"); and third, they lack an LLM content reorganization strategy tailored to the construction domain, failing to automatically construct differentiated prompts based on rule action types (extraction and reorganization, data mining, knowledge base fusion, or flowchart generation).

[0027] This invention proposes an automatic generation method for engineering construction plan documents based on a rule engine. The method first receives the project name, domain, and template path input by the user, loads preset template rules (including rule identifiers and action types), and indexes the template document and project data separately. Then, it merges these two to construct a full-text search engine index based on Chinese binary word segmentation technology, enabling the automatic extraction of key project information (project name, participating units, construction period, etc.). Based on this, a topological sort is performed according to the directed acyclic graph dependencies between template rules, and various rule actions are executed sequentially—including copying and replacing parameters from the knowledge base template, filling variables according to fixed sentence structures, searching project data via FTSS and extracting and recombining it using LLM, keyword retrieval and LLM information mining, LLM fusion generation of the knowledge base framework and project content, and extracting procedures from upstream chapters and generating flowchart descriptions using LLM—thus generating the content of each chapter in an orderly manner. Finally, a Word document retaining the cover, style, and header / footer is constructed based on the template document. After cleaning template placeholders, an edited version containing source tags and specification comparison results, and a clean version with annotations removed, are simultaneously saved, achieving dual-version delivery.

[0028] This invention represents a significant advancement over existing technologies: First, it dramatically improves compilation efficiency, reducing the time required for traditional scheme compilation from days or weeks to minutes. Furthermore, its rule-based, dependency-driven execution system covers all core chapters from project overview to emergency response, achieving a scheme completeness of over 85%, with only 15% of project-specific data requiring manual verification. Second, it effectively ensures the accuracy and consistency of standard references by automatically comparing versions from the knowledge base, construction organization, and template, avoiding omissions and errors from manual verification. Third, the unique cross-chapter information transmission relationships in the engineering field (such as previous outputs automatically becoming the context for subsequent inputs) are precisely maintained through the DAG execution engine, resolving the issue of contextual breaks in construction process descriptions. Finally, the dual-version output mechanism balances internal quality control with external delivery needs—the edited version retains the source markers and standard comparison results for each piece of content, facilitating auditing and traceability, while the clean version removes all markers and can be directly used for formal delivery. These two versions meet different usage scenarios, comprehensively improving the intelligence and automation level of engineering construction scheme compilation.

[0029] To make the objectives, technical solutions, and advantages of this application clearer, the embodiments of this application will be described in further detail below with reference to the accompanying drawings.

[0030] Example 1 Firstly, this invention provides a method for automatically generating engineering construction plan documents based on a rule engine, the flowchart of which is shown below. Figure 1 As shown, it includes the following steps: Receive user input for the project name, project domain, and template path; In response to the received project name, project domain, and template path, a preset template rule is loaded, wherein the template rule includes at least a rule identifier and a rule action type; An index is created for the template documents specified by the template path, generating a template index; Load and index project data, and generate a project data index; Based on the template index and the project data index, a full-text search engine index is constructed; The key information of the project is extracted from the project data according to the full-text search engine index. The key information of the project includes at least the project name, participating units and construction period information. Based on the directed acyclic graph dependency order among the template rules, the template rules are executed sequentially to generate the content of each chapter based on the key project information; Based on the template index, the content of each chapter is filled into the template document to construct an editable Word document, which includes content source tags and specification comparison results; Remove the comment marks from the edited version of the Word document to generate a clean version of the Word document; Output the edited version of the Word document and the clean version of the Word document.

[0031] One specific flowchart of a rule engine-based method for automatically generating engineering construction plan documents is shown below. Figure 2 As shown.

[0032] Furthermore, the rule action type includes at least one of the following: Copy and replace parameters from knowledge base template (the variable "kb_copy_and_replace" represents the action type of copying and replacing parameters from knowledge base template), used to copy chapter content from knowledge base template and perform parameter replacement and specification comparison; Fill variables according to a fixed sentence pattern (the variable "template_fill" represents the action type of filling variables according to a fixed sentence pattern), which is used to fill variables according to a fixed sentence pattern template; The method involves searching for project information using FTSS and extracting and reconstructing it using LLM (the variable "extract_and_rephrase" represents the action type of searching for project information using FTSS and extracting and reconstructing it using LLM). This method is used to search for project information using a full-text search engine and extract and reconstruct it using a large language model. Keyword retrieval and LLM information mining (the variable "project_docs_extract" is defined to represent the action type of keyword retrieval and LLM information mining) are used to perform keyword search through the full-text search engine and extract information using a large language model; The knowledge base framework and project content are integrated into an LLM model for generation (the variable "kb_adapt_and_fill" represents the action type of the LLM integration of the knowledge base framework and project content), which is used to integrate and generate based on the knowledge base framework and project content using a large language model. Extracting process steps from upstream chapters and generating flowchart descriptions using LLM (defining the variable "flowchart_generate" to represent the action type for extracting process steps from upstream chapters and generating flowchart descriptions using LLM), is used to extract process information from upstream chapters and generate flowchart descriptions using a large language model.

[0033] Furthermore, the construction of the full-text search engine index includes: Based on the in-memory dictionary index, it categorizes and retrieves data by document type. It parses the heading hierarchy structure of .docx files (defining the variable "Heading" as the heading hierarchy structure), creating structured records for the headings and body text of each chapter. It supports categorized retrieval by document type (construction organization / geological exploration / design) and provides a field extraction method (defining the variable "extract_field" to represent the field extraction method) to extract specific information by field name (e.g., "project name", "construction unit"). The field extraction method is a function method for extracting data.

[0034] Import all document content into an SQLite FTS5 virtual table to enable full-text search based on an inverted index, supporting Boolean operations, phrase matching, and result segment highlighting.

[0035] The template document and project data are segmented using Chinese bigram segmentation technology, and the full-text search engine index is constructed based on the segmentation results. For example, Chinese text is segmented into two adjacent characters by Chinese bigram segmentation (referred to as Slide WindowBigram in English) using a Chinese bigram segmentation method (defined by the variable "chinese_tokenize" to represent the Chinese bigram segmentation method), which solves the problem of poor support for Chinese by the default unicode61 segmenter of FTS5.

[0036] Multi-keyword combination search: The multi-keyword combination search method (defined by the variable "search_multi") searches for each keyword individually and merges them to remove duplicates, avoiding the omission of single keywords.

[0037] Furthermore, the template rules are executed sequentially according to the directed acyclic graph dependency order among the template rules, including: The dependencies between template rules are analyzed, and a directed acyclic graph (DAG) is constructed. Dependencies are defined by pre-defining the dependencies between chapters using a strict dependency dictionary (defined by the variable "strict_deps"). For example, "4_4 (Main Construction Methods)" depends on "4_3 (Process Flow)" and "5_3 (Quality Technical Measures)" depends on "4_4 (Main Construction Methods)". The strict dependency dictionary is a Python dictionary with the current rule as the key and its list of dependent rules as the value.

[0038] The execution order of each template rule is determined according to the topological sorting of the directed acyclic graph. Topological sorting execution means: a preset dependency order list (defined by the variable "dependency_order") represents the sequential traversal of rules. For each rule, it is checked whether its dependent rules have been successfully executed. If not, the rule is skipped and recorded. Each template rule is executed sequentially according to the execution order. After the current rule is completed, the execution of its downstream dependent rules is triggered.

[0039] The `dependency_order` list is a Python list containing unique identifiers (IDs) of all rules to be executed. It determines the order in which the rule executor attempts to execute rules. The `dependency_order` list is fundamentally different from a strict dependency dictionary (defined by the variable "strict_deps"). Their core difference lies in their roles: strict_deps defines the logical threshold of "whether it can be done", while dependency_order defines the traversal order of "which to try first".

[0040] Specifically, `strict_deps` is a set of hard rule dependency locks. It clearly declares the logical prerequisites between sections through a dictionary structure; for example, "Main Construction Methods" can only be executed after "Process Flow" is successful. This is a purely business logic constraint that focuses on the correctness and completeness of content generation, ensuring that downstream sections are not generated rashly when upstream data is missing, thereby avoiding logically incorrect content.

[0041] The `dependency_order` is a flexible sequence of scheduling attempts. It's simply a list arranged in the natural chapter order of a document, determining the order in which the executor iterates through the documents. It doesn't carry any logical information about "who it must wait for"; it's more like a "to-do list" for the system to check sequentially. Even if a chapter's dependency hasn't been satisfied, the executor will still process it, but it will skip it if the dependency check fails.

[0042] The way the two work together also reflects this division of labor: the executor processes rules one by one strictly according to the dependency_order, but before processing each rule, it first uses strict_deps to check whether all its predecessor dependencies have been successfully verified. This "sequential traversal + on-demand locking" mechanism makes the method execution both clear in its execution path and subject to strict logical constraints.

[0043] Furthermore, the `dependency_order` table design offers a significant advantage: greater fault tolerance and flexibility. In strict topology sorting, the entire subsequent process will stall if an upstream node fails. However, in the method of this invention, even if a dependency is temporarily not satisfied, the executor will only record and skip the current rule, continuing to process subsequent irrelevant sections. This allows the system to continue working as much as possible even with missing data, and leaves room for subsequent "retries" or "supplementary execution," making the entire document generation process more robust.

[0044] This invention employs a flexible strategy of "traversal + skipping" rather than static sorting: Strict topological sort (A→B→C): Must be executed linearly, and multiple retry rounds are not allowed.

[0045] The strategy of this invention (list traversal + skip): even if it is 5_3's turn, if the dependency 4_4 fails, skip it; when it is 5_3's turn again in the next loop (or the next round of scheduling), if 4_4 has succeeded, then it can be executed.

[0046] Therefore, the `dependency_order` list is usually arranged directly according to the natural chapter order of the document (e.g., 1.1, 1.2...4.3, 4.4...8.1). This design has two advantages: 1. Intuitive: The code logic is consistent with the document's writing order, making it easier to maintain. 2. Strong fault tolerance: Even if a previous rule fails, the system will not freeze, but will record the failure and continue trying subsequent chapters; if the previous rule succeeds later, the dependency will be automatically unblocked.

[0047] Furthermore, the extraction of key project information from the project data includes: Use a full-text search engine to retrieve keywords related to the project name, participating units, and construction period from the project data; By using a large language model to extract information from the search results, structured key project information is obtained.

[0048] Furthermore, based on the template index, the content of each chapter is filled into the template document to construct an editable Word document, specifically including: Fill the corresponding positions with the content of each chapter according to the chapter structure of the template document; The original cover format, style definition, header and footer settings of the template document are retained.

[0049] After building the editable Word document, the process also includes: The edited version of the Word document retains source tags, which are used to indicate the original source of each chapter's content; The standard comparison results are retained in the edited version of the Word document. These results are used to indicate the conformity between the generated chapter content and the preset standards.

[0050] Furthermore, the step of removing comment markers from the edited version of the Word document also includes: Identify template placeholder marks in the template document, including placeholders of the "XX Project" type; Replace the template placeholder with the corresponding actual information in the project's key information.

[0051] Furthermore, it also includes a three-level specification version comparison method to determine whether the generated Word document meets the requirements of the template rules. The specific steps include: Extract the standard numbers from the template rules, use regular expressions to match the preset set of standard prefixes, and extract all referenced standard standard numbers from the template document; The first level of comparison is knowledge base comparison, which involves performing cardinality matching between the extracted specification number and the specification files stored in the specification knowledge base. The cardinality matching ignores the year and version number differences in the specification number and only compares the main body number part of the specification. If there is any version of the specification file with the same main body number as the current specification in the knowledge base, the latest version information of the specification in the knowledge base and the current template reference version information are recorded, and it is determined that the knowledge base has been matched. The second level of comparison is the construction organization comparison. For specifications that fail to match in the knowledge base after the first level of comparison, the referenced specification number is extracted from the project construction organization design document. The unmatched specification is compared with the specification number extracted from the construction organization design document. If there is a specification reference record in the construction organization that has the same main specification number as the current specification, the version information referenced in the construction organization is recorded and it is determined that the construction organization has been matched. The third level of comparison is to mark the specification as questionable. For specifications that are not matched after the first and second levels of comparison, they are marked as questionable because none of the three parties have this specification, and the questionable mark is embedded in the generated document.

[0052] Example 2 Secondly, this application provides an automatic generation system for engineering construction plan documents based on a rule engine, including: The input module is used to receive user input of the project name, project domain, and template path; The data processing module is used to load preset template rules in response to the received project name, project domain and template path. The template rules include at least a rule identifier and a rule action type. An index is created for the template documents specified by the template path, generating a template index; Load and index project data, and generate a project data index; Based on the template index and the project data index, a full-text search engine index is constructed; Extract key project information from the project data. The key project information includes at least the project name, participating units, and construction period. Based on the directed acyclic graph dependency order among the template rules, the template rules are executed sequentially to generate the content of each chapter based on the key project information; The output module is used to fill the content of each chapter into the template document based on the template index, and construct an editable Word document, wherein the editable Word document contains content source tags and specification comparison results; Remove the comment marks from the edited version of the Word document to generate a clean version of the Word document; Output the edited version of the Word document and the clean version of the Word document.

[0053] Furthermore, an embodiment is provided, showing a system diagram for automatically generating engineering construction plan documents based on a rule engine, as follows: Figure 3 As shown, an automatic generation system for engineering construction plan documents based on a rule engine includes: The web control panel provides a unified entry point for multiple workflows, and the backend routing layer is built on the FastAPI framework to receive the project name, project domain and template path input by the user and route the request to the project management module. The project management module communicates with the backend routing layer and is used to create project instances and maintain project context information. The rules engine, connected to the project management module, serves as a unified entry point for document generation. The rules engine integrates at least the following: Template indexer, used to create a template index for standardized scheme template documents under a specified template path; A document indexer is used to load and index project materials, including project specifications and functional design documents. The rule executor has N preset domain rules (preferably 47), and there are directed acyclic graph (DAG) dependencies between the domain rules. The rule executor executes each rule in sequence according to the topological order of the DAG dependencies to generate the content of each chapter. The storage layer includes at least: A standardized solution template library for storing standardized solution templates in .docx format; A knowledge base is used to store historical standardization schemes and a library of specifications and standards. The project database is used to store the specification documents and functional design data of the current project. The index layer includes: SQLite Full-Text Search Engine (FTSS) uses Chinese ternary word segmentation technology to create a full-text index of the templates, knowledge base content and project materials. An in-memory dictionary index is used to cache key project information and intermediate variables during rule execution. The content processing layer includes: The parameter replacer is used to replace placeholders in the template with actual project parameters; The LLM reorganization pipeline communicates with an external large language model API (preferably DeepSeek) to construct differentiated prompt words based on different rule action types, and calls the large language model for content extraction, reorganization, and fusion generation; The Level 3 specification comparison module is used to perform a three-way cross-comparison of the specification version in the knowledge base, the specification version referenced in the project materials, and the specification version built into the template, and output the specification comparison results. The document building and output layer includes: A document builder is used to construct a Word document based on the standardized template and the content of each chapter, while retaining the cover, style, header and footer of the template. The dual-version output module is used to generate and output an edited version of the Word document and a clean version of the Word document. The edited version retains the content source tags and the standard comparison results, while the clean version removes all annotation tags and source tags.

[0054] Innovation Point 1: Template Chapter - Automatic Rule Matching Engine The template indexer of this invention (implemented by calling template_indexer.py) achieves automatic mapping from unstructured Word templates to structured rules. The specific steps are as follows: S1 Load Domain Rule Set: Load the corresponding chapter rule dictionary (i.e., SECTION_RULES dictionary) according to the engineering domain (port and waterway / road and bridge / municipal / water conservancy / construction / mechanical and electrical / rail transit / environmental protection). Each rule defines the chapter number, title keyword pattern (defined by the variable "title_patterns"), execution action type (defined by the variable "action"), parameter source (defined by the variable "source"), etc.

[0055] S2 parses the template document structure: Iterates through the XML document body elements of the .docx document, identifies the heading level of each paragraph by the Heading style or outlineLvl attribute, and extracts the body paragraph text and table data.

[0056] S3 two-layer matching strategy: Precise number matching: Extract the number prefix in the title (e.g., "2.1.1") and compare it precisely with the underscore number in the rule name (e.g., "2_1_1"). If a match is found, it is directly mapped.

[0057] Keyword fuzzy matching: When a number match fails, the keyword hit count is calculated between the title keyword pattern (title_patterns) and the title text. The match with the highest hit count (≥ minimum hit count (min_hits) and the highest score is selected. Hitting a number prefix results in a significant weighting (+10 points).

[0058] S4 Smart Subheading Merge: When a subheading matches the same rule, the content is automatically merged into its corresponding parent heading (i.e., a higher-level heading) to avoid content truncation.

[0059] Innovation Point Two: Rule Execution Scheduler for DAG Dependency Resolution The rule executor of this invention (implemented by calling rule_executor.py) manages the execution order of 47 rules based on a directed acyclic graph (DAG): Dependency definition: Dependencies between chapters are predefined through a strict dependency dictionary, such as "4_4 (Main construction methods)" depending on "4_3 (Process flow)" and "5_3 (Quality and technical measures)" depending on "4_4 (Main construction methods)".

[0060] Topological sorting execution: Traverse the rules in the order of the preset dependency order list, check whether the dependent rules of each rule have been successfully executed, skip the rule if not, and record the result.

[0061] The six motion processors are shown in Table 1: Table 1. List of six motion processor processing schemes

[0062] Multi-template fallback search: When a chapter is not matched in the format template, the search is automatically performed in the content reference template list. If no match is found, fuzzy matching is performed (the variable "fuzzy_match" is defined to represent fuzzy matching).

[0063] Innovation Point 3: Three-Level Specification Version Comparison System The parameter replacer of this invention is implemented through the "ParamReplacer" module, achieving the purpose of three-level cross-validation of engineering specifications: S1 Extract the specification number from the template: Use regular expressions to match the standard specification number with prefixes such as JTS / JTG / GB / JGJ / SL / TB / CJJ / CECS / AQ / DL.

[0064] S2 Level 1—Knowledge Base Comparison: Search for the specification number in all filenames in the specification knowledge base directory and perform cardinality matching with the template specification (ignoring differences in year and version number). For example, “GB 50204-2015” and “GB 50204-2002” are considered different versions of the same specification.

[0065] S3 Level 2 - Construction Organization Comparison: Extract the referenced specification numbers from the project construction organization design documents and compare them with specifications not found in the knowledge base for confirmation.

[0066] S4 Level 3 - Marked as Questionable: Specifications that have not been matched in the first two levels are marked as "No such specification exists in any of the three parties".

[0067] Radix matching algorithm: It extracts the prefix letter + number part of the specification by using the standard basic matching method (the variable "spec_base_match" is defined to represent the standard basic matching method) and compares them, ignoring the version year difference, and solves the problem of automatic judgment when multiple versions of the same specification coexist.

[0068] Innovation Point 4: LLM Differentiated Reconstituted Pipelines for the Construction Industry The LLM pipeline of this invention (implemented by calling "llm_pipeline.py") designs four differentiated prompt word strategies for different rule action types: (1) Extract and Rephrase (define the variable "extract_and_rephrase" to indicate extraction and rephrase): construct a context containing "rule description + source requirements + relevant paragraphs found", instruct LLM to rephrase into 1-3 paragraphs, retain key data and remove irrelevant details.

[0069] (2) Project data mining (defined as “project_docs_extract” to represent project data mining): Search all data types (not limited to implementation organization), instruct LLM to organize according to the scheme format (retain the number skeleton), extract data directly from the original text without fabrication, and mark missing items as “requires manual supplementation”.

[0070] (3) Knowledge base adaptation and integration (defined as “kb_adapt_and_fill” to indicate knowledge base adaptation and integration): It provides both the knowledge base (KB) template framework structure and the construction operation content of this project. The instruction big language model (LLM) retains the three-level heading structure but replaces it with the specific parameters of the project, and adds project-specific measures.

[0071] (4) Flowchart generation (defined as “flowchart_generate” to represent flowchart generation): Search for construction procedures from the construction organization, combine the content of the upstream chapters, and instruct the large language model to output 7-15 arrow-shaped process flow, and insert diamond judgment nodes after the inspection / acceptance nodes.

[0072] API configuration self-discovery: The system automatically reads API keys and endpoints from Hermes configuration files or environment variables without hardcoding.

[0073] Innovation Point 5: FTS5 Chinese Full-Text Index and Field Extraction Dual-Channel Retrieval The document indexer (doc_indexer.py) of this invention implements a dual retrieval mechanism: S1 Memory Dictionary Index: Parses the Heading structure of .docx files, establishes the title and body of each chapter as a structured record, supports categorized retrieval by document type (construction organization / geological exploration / design), and provides methods to extract specific information by field name (such as "project name" or "construction unit").

[0074] S2 SQLite FTS5 Full-Text Index: Imports all document content into an SQLite FTS5 virtual table to achieve efficient full-text search based on an inverted index. It supports Boolean operations, phrase matching, and snippet functions. In FTS5, snippet() is a built-in helper function used to extract contextual fragments containing keywords from matched documents and mark them as highlighted.

[0075] S3 Chinese Bigram Segmentation: This feature segments Chinese text into pairs of adjacent characters using a Chinese word segmentation method (Slide Window Bigram), resolving the issue of poor Chinese support from the default unicode61 word segmenter in FTS5.

[0076] S4 Multi-keyword combination search: This method searches for each keyword individually and merges them to remove duplicates, avoiding the omission of single keywords.

[0077] Compared with the prior art, the present invention has the following beneficial effects: The efficiency of the preparation has been greatly improved: the preparation time of the construction plan has been reduced from several days to several minutes. 47 rules are automatically executed in the order of dependency, covering 11 chapters from project overview to emergency response, and the completeness of the generated plan can reach more than 85% (the remaining 15% is project-specific data that needs to be manually confirmed).

[0078] Ensuring the accuracy of standard references: The three-level standard version comparison automatically detects obsolete / expired standards referenced in the template, and obtains the latest version information from the knowledge base and construction organization, avoiding omissions in manual standard verification.

[0079] It has broad applicability: It supports eight major engineering fields, including ports and waterways, roads and bridges, municipal works, water conservancy, construction, electromechanical engineering, rail transit, and environmental protection. Each field is equipped with an independent set of 47 rules, which can be flexibly expanded to new fields.

[0080] Quality traceability: Dual-version output mechanism - the edited version retains the source tags (KB template / fixed sentence / LLM reorganization) and specification comparison results for each piece of content, while the clean version removes all tags and can be delivered directly to meet different use cases.

[0081] The system is highly scalable: the web control panel adopts a plug-in architecture, and the six action handlers are registered through the action handler registration dictionary (referred to as ACTION_HANDLERS dictionary). To add a new action type, only a new handler function needs to be added, without modifying the core scheduling logic.

[0082] Example 3 Example 3 is based on Examples 1 and 2, and takes the automatic generation of bridge bored pile construction scheme as an example for illustration.

[0083] Step 1: Project Creation and Context Initialization The user enters project information in the Web control panel: the project name is "Chang'an Lock Bridge Bored Pile", the field is "road and bridge", and the plan type is "special plan". The system creates draft / and output / subdirectories under the project root directory, and writes a context configuration file (namely the context.json file) containing domain marking rules, supplementary prompts and knowledge base template matching results.

[0084] Step 2: Template Indexing The system loads "JS080301010 Bridge Bored Pile Construction Plan Template.docx", and parses the template structure through the template indexing method (namely the index_template_v2 method). For the title "2.1.1 Project Introduction", its numeric prefix "2.1.1" exactly matches the rule number "2_1_1", and the mapped execution action type is "extract and rephrase" (corresponding to extract_and_rephrase); for the title "4.3 Process Flow", it matches the rule number "4_3", whose execution action type is "flowchart generation" (corresponding to flowchart_generate), and marks its dependent rule as "3_1" (that is, the depends_on value is "3_1").

[0085] Step 3: Project Data Indexing The system scans the construction organization.docx file under the project file directory (namely the project_files / construction_organization / path), establishes an in-memory dictionary index and a FTS5 full-text index through the document indexer (namely the DocIndexer class), and extracts the content of each chapter from the heading style (namely Heading style) structure. It calls the field extraction method (namely the extract_field method), passes in the parameters "project name" and "construction organization" to obtain the official name of the project, and extracts information such as the construction owner, construction unit, and supervision unit.

[0086] Step 4: Sequential Execution of Rules Traverse 47 rules according to the dependency_order list: • Rule 2_1_1 (Basic Project Information) → extract and rephrase method (namely extract_and_rephrase) → search keywords such as "project overview / project introduction / project scale" through FTS5 full-text search (namely FTS5) → reorganized into 1 to 3 paragraphs by a large language model (LLM).

[0087] • Rule 4_3 (Process Flow) → Flowchart generation method (i.e., flowchart_generate) → Check that dependency 3_1 has been successfully executed → Search for "construction process flow / construction procedure" through FTS5 full-text search → Output the process "construction preparation → surveying and setting out → casing installation → drilling rig positioning → drilling and hole formation → hole cleaning → reinforcement cage installation → concrete pouring → pile testing" from the Large Language Model (LLM).

[0088] • Rule 4_4 (Main Construction Method) → Knowledge Base Copy and Replace Method (i.e., kb_copy_and_replace) → Check that dependency 4_3 has been successfully executed → Copy method description from template → Perform parameter replacement.

[0089] • Rule 6_2 (Hazard Identification) → Knowledge Base Copy and Replace Method (i.e., kb_copy_and_replace) → Copy the hazard table from the template.

[0090] • Rule 6_4 (Security Technical Measures) → Knowledge Base Copy and Replace Method (i.e., kb_copy_and_replace) → Check that dependency 6_2 has been successfully executed → Replace parameters based on template content.

[0091] • Rule 9_1 (Acceptance Criteria) → Knowledge Base Copy and Replace Method (i.e., kb_copy_and_replace) → Trigger Level 3 Specification Comparison → Detected that the specification number “GB 50204-2002” has been updated to “GB 50204-2015” in the knowledge base → Marked as updated.

[0092] Step 5: Documentation Building and Dual Version Output The system opens the format template.docx, retaining the cover (including the approval unit table), headers, footers, and body text styles. It deletes the template description pages (such as "Scheme Template Compilation Instructions" and "Instructions for Use"), replaces "XX Project" in the cover with the actual project name, and replaces the placeholders for participating units in Section 2.6 with the actual unit names. It outputs the generated content rule by rule, inserting source tags and specification comparison results. After saving the edited version, it removes all tag comments using the comment removal method (i.e., the strip_annotations method), generating a clean version.

[0093] Output: Edited version .docx (including marks such as "Source: Knowledge Base Template", "Source: Large Language Model Reorganization", "Standard JTS 257-2008 → JTS 257-2008 Confirmed") + Clean version .docx (formatted solution document that can be delivered directly).

[0094] The above-described embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application, and should all be included within the protection scope of this application.

Claims

1. A method for automatically generating engineering construction plan documents based on a rule engine, characterized in that, Includes the following steps: Receive user input for the project name, project domain, and template path; In response to the received project name, project domain, and template path, a preset template rule is loaded, wherein the template rule includes at least a rule identifier and a rule action type; An index is created for the template documents specified by the template path, generating a template index; Load and index project data, and generate a project data index; Based on the template index and the project data index, a full-text search engine index is constructed; The key information of the project is extracted from the project data according to the full-text search engine index. The key information of the project includes at least the project name, participating units and construction period information. Based on the directed acyclic graph dependency order among the template rules, the template rules are executed sequentially to generate the content of each chapter based on the key project information; Based on the template index, the content of each chapter is filled into the template document to construct an editable Word document, which includes content source tags and specification comparison results; Remove the comment marks from the edited version of the Word document to generate a clean version of the Word document; Output the edited version of the Word document and the clean version of the Word document.

2. The method for automatically generating engineering construction plan documents based on a rule engine according to claim 1, characterized in that, The rule action type includes at least one of the following: Knowledge base copy and populate is used to copy chapter content from a knowledge base template and perform parameter replacement and specification comparison. Template variable filling is used to fill variables according to a fixed sentence template. Intelligent extraction and rewriting is used to search for project information through a full-text search engine and extract and reorganize it using a large language model; Item-by-item data extraction is used for keyword searching through a full-text search engine and information extraction using a large language model; Framework fusion filling is used to generate data by fusing knowledge base frameworks and project content using a large language model. Intelligent process generation is used to extract process information from upstream chapters and generate flowchart descriptions using a large language model.

3. The method for automatically generating engineering construction plan documents based on a rule engine according to claim 1, characterized in that, The construction of the full-text search engine index includes: The template document and the project information are segmented using Chinese binary word segmentation technology, and the full-text search engine index is constructed based on the word segmentation results. Import all document content into an SQLite FTS5 virtual table to enable full-text search based on an inverted index; Based on the in-memory dictionary index, retrieve documents categorized by type; Multi-keyword combination search, searching by keyword one by one and merging to remove duplicates.

4. The method for automatically generating engineering construction plan documents based on a rule engine according to claim 1, characterized in that, Based on the directed acyclic graph dependency order among the template rules, the template rules are executed sequentially, including: Analyze the dependencies between the template rules and construct a directed acyclic graph; The execution order of each template rule is determined according to the topological sorting of the directed acyclic graph. Each template rule is executed sequentially according to the execution order described above. Once the current rule is completed, the execution of its downstream dependent rules is triggered.

5. The method for automatically generating engineering construction plan documents based on a rule engine according to claim 1, characterized in that, Extract key project information from the project data, including: Use a full-text search engine to retrieve keywords related to the project name, participating units, and construction period from the project data; By using a large language model to extract information from the search results, structured key project information is obtained.

6. The method for automatically generating engineering construction plan documents based on a rule engine according to claim 1, characterized in that, Based on the template index, the content of each chapter is filled into the template document to construct an editable Word document, specifically including: Fill the corresponding positions with the content of each chapter according to the chapter structure of the template document; The original cover format, style definition, header and footer settings of the template document are retained.

7. The method for automatically generating engineering construction plan documents based on a rule engine according to claim 1, characterized in that, After building the editable Word document, the process also includes: The edited version of the Word document retains source tags, which are used to indicate the original source of each chapter's content; The standard comparison results are retained in the edited version of the Word document. These results are used to indicate the conformity between the generated chapter content and the preset standards.

8. The method for automatically generating engineering construction plan documents based on a rule engine according to claim 1, characterized in that, The process of removing comment markers from the edited Word document also includes: Identify template placeholder marks in the template document; Replace the template placeholder with the corresponding actual information in the project's key information.

9. The method for automatically generating engineering construction plan documents based on a rule engine according to claim 1, characterized in that, It also includes a three-level specification version comparison method to determine whether the generated Word document meets the template rules. The specific steps include: Extract the standard numbers from the template rules, use regular expressions to match the preset set of standard prefixes, and extract all referenced standard standard numbers from the template document; The first level of comparison is knowledge base comparison, which involves performing cardinality matching between the extracted specification number and the specification files stored in the specification knowledge base. The cardinality matching ignores the year and version number differences in the specification number and only compares the main body number part of the specification. If there is any version of the specification file with the same main body number as the current specification in the knowledge base, the latest version information of the specification in the knowledge base and the current template reference version information are recorded, and it is determined that the knowledge base has been matched. The second level of comparison is the construction organization comparison. For specifications that fail to match in the knowledge base after the first level of comparison, the referenced specification number is extracted from the project construction organization design document. The specification that fails to match in the knowledge base is compared with the specification number extracted from the construction organization design document. If there is a specification reference record in the construction organization that has the same main specification number as the current specification, the version information referenced in the construction organization is recorded and it is determined that the construction organization has been matched. The third level of comparison is to mark the specification as questionable. For specifications that are not matched after the first and second levels of comparison, they are marked as questionable because none of the three parties have this specification, and the questionable mark is embedded in the generated document.

10. A rule-based engine-based automatic generation system for engineering construction plan documents, characterized in that, include: The input module is used to receive user input of the project name, project domain, and template path; The data processing module is used to load preset template rules in response to the received project name, project domain and template path. The template rules include at least a rule identifier and a rule action type. An index is created for the template documents specified by the template path, generating a template index; Load and index project data, and generate a project data index; Based on the template index and the project data index, a full-text search engine index is constructed; Extract key project information from the project data. The key project information includes at least the project name, participating units, and construction period. Based on the directed acyclic graph dependency order among the template rules, the template rules are executed sequentially to generate the content of each chapter based on the key project information; The output module is used to fill the content of each chapter into the template document based on the template index, and construct an editable Word document, wherein the editable Word document contains content source tags and specification comparison results; Remove the comment marks from the edited version of the Word document to generate a clean version of the Word document; Output the edited version of the Word document and the clean version of the Word document.

Citation Information

Patent Citations

  • Comprehensive pipeline space-time analysis and risk estimation method based on 4D-BIM technology

    CN114418330A

  • System and method for automated document generation

    US11087219B1

  • Template system for custom document generation

    US20150046791A1