Engineering consulting report ai agent generation and signability review method

CN122817480APending Publication Date: 2026-09-25YINSHU TECHNOLOGY (CHENGDU) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202611000825.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-07
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

[0009]针对现有技术的不足,本发明提供了工程咨询报告AI智能体生成与可签发性审查方法,解决了现有利用人工智能生成工程咨询报告的方法难以处理工程参数间复杂的逻辑依赖关系,且模型常凭空生成无事实依据的项目数据;此外,由于缺乏针对数值逻辑、法规时效与文档格式的严格审查机制,导致生成的报告准确度低,无法直接满足工程咨询行业的签发要求的问题

Benefits of technology

[0085]1、本发明通过引入必须实体规则与不可编造实体清单,限定了确定性信息闭包。在生成报告阶段,系统拦截模型对缺乏数据源的项目参数的生成行为,并使用结构化占位符进行替代。该机制将文本语句生成与事实数据赋值进行物理解耦,避免了模型在生成专业报告时伪造底层工程数据的现象,确保了报告中核心事实指标的真实性。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122817480A_ABST
    Figure CN122817480A_ABST
Patent Text Reader

Abstract

The application relates to the field of artificial intelligence and engineering digital information technology, and discloses an engineering consultation report AI intelligent body generation and signability examination method, which comprises the following steps: analyzing engineering data to extract multidimensional data items, calling corresponding entity constraint rules, searching reference content from an expert report knowledge base, and shielding project-specific factual information; when generating a preliminary draft, a deterministic information closure is constructed, it is judged whether a target entity belongs to the closure, if not, generation of a sourceless value is prevented and a structured placeholder is inserted; after receiving user supplementary data, derivative nodes are calculated by using an engineering data dependency graph, the placeholder is replaced, and only the relevant area is executed to locally rewrite the report text; finally, the rewritten text is executed to automatically examine numerical consistency, boundary threshold and regulation time limit, and a signability examination result is output. The application avoids fabricating underlying engineering data when a model generates a report, and guarantees logical synchronization of associated derivative data and compliance of report content.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of artificial intelligence and engineering digital information technology, specifically to a method for generating AI intelligent agents and reviewing the issueability of engineering consulting reports. Background Technology

[0002] The preparation of engineering consulting reports involves the analysis of engineering data, parameter calculation, assessment of the applicability of laws and standards, and collaborative updating of multiple chapters. With the development of large language models and knowledge base technologies, utilizing artificial intelligence to assist in the generation and review of professional reports has become the development direction of engineering digitalization.

[0003] Chinese invention patent application CN122113890A discloses a text generation method based on a large language model. This method generates a structured outline and citation constraint prompts based on text generation instructions and reference materials, calls the large language model to generate initial text with citation identifiers, and outputs the final text through source consistency verification, external information verification, and text correction.

[0004] The methods described above can improve the reliability and traceability of the generated text, but they mainly focus on verifying the correspondence between the generated statements and the reference materials. For project entities in engineering consulting reports, such as client names, project locations, equipment parameters, engineering drawings, site photos, and investment amounts, which cannot be determined by the model itself, there is still a lack of pre-generation control mechanisms based on entity type and source status. When engineering data is missing, it is difficult to prevent entities with no definite source from being included in the report text, and it is also difficult to establish a correspondence between missing entities, report locations, and subsequent supplementary data using unique identifiers.

[0005] Chinese invention patent application CN121981085A discloses an intelligent identification and verification method for engineering technical standards based on multi-source information fusion. This method cleanses and extracts text from engineering documents, identifies engineering technical standards within the documents through a multi-path collaborative approach, and performs deduplication, comparison, and missing item analysis on the standard list.

[0006] The above methods can help identify and verify technical standards in engineering documents, but they mainly deal with existing standard reference information in the documents, and are difficult to solve the problems of controlling the source of engineering entities, filling in missing data, and backfilling supplementary data in the process of generating engineering consulting reports.

[0007] Furthermore, parameters in engineering consulting reports, such as production scale, raw material consumption, pollutant generation, treatment efficiency, and emissions, often exhibit calculation dependencies. When upstream parameters are supplemented or modified, downstream indicators need to be calculated simultaneously and the relevant report paragraphs updated. Existing technologies struggle to identify affected data nodes and text areas based on the dependencies between engineering parameters, and also find it difficult to simultaneously conduct multi-dimensional reviews, including numerical consistency, boundary thresholds, regulatory timeliness, and missing entities.

[0008] Therefore, it is necessary to provide an AI-powered method for generating and issuing engineering consulting reports, establishing unified identifiers and source traceability relationships for engineering entities, inserting structured placeholders when an entity lacks a definite source, calculating derived entities based on engineering data dependencies and partially updating relevant text after acquiring supplementary or changed data, and automatically reviewing the data consistency, indicator rationality, and regulatory and standard status in the report. Summary of the Invention

[0009] To address the shortcomings of existing technologies, this invention provides a method for generating and reviewing the issueability of engineering consulting reports using AI intelligent agents. This method solves the problems of existing methods for generating engineering consulting reports using artificial intelligence struggling to handle complex logical dependencies between engineering parameters, and the frequent generation of project data without factual basis by the models. Furthermore, the lack of a strict review mechanism for numerical logic, regulatory timeliness, and document format leads to low accuracy in the generated reports, which cannot directly meet the issuance requirements of the engineering consulting industry.

[0010] To address the above problems, the present invention provides the following technical solution:

[0011] This invention provides a method for generating AI agents and reviewing the issueability of engineering consulting reports, including:

[0012] Obtain a vertical AI agent for engineering consulting reports based on historical engineering consulting reports that have been issued or reviewed by experts, and build an expert report knowledge base.

[0013] Parse the engineering data submitted by the user, extract multi-dimensional data items, determine the target report type, and retrieve the corresponding entity rule file and list of non-fabricable entities;

[0014] Based on the multidimensional data items, reference content is retrieved from the expert report knowledge base, and item-specific factual information is masked.

[0015] When generating the initial draft of the engineering report, a deterministic information closure is constructed for the target entity objects corresponding to the entity rule file and the list of non-fabricable entities. The deterministic information closure consists of the multidimensional data items, supplementary data confirmed by the user, fixed threshold data in the regulatory standard library, and result data calculated by the system based on deterministic formulas.

[0016] Determine whether the target entity object belongs to the deterministic information closure. If not, prevent the generation of passive value description and insert a structured placeholder.

[0017] Receive supplementary data from the user, replace the structured placeholders, and generate a rewrite report text;

[0018] The rewritten report text is automatically reviewed, and the results of the approval review are output.

[0019] This invention decouples semantic logic generation and factual data assignment in engineering reports by constructing entity rules and a list of non-fabricable entities. During report generation, the system restricts the model's authority to generate parameters for non-source projects by defining deterministic information closures and uses structured placeholders for substitution, preventing data fabrication. In the data supplementation stage, topological sorting based on a directed acyclic graph ensures the synchronous conversion and updating of derived data, and automatic review logic performs data conflict comparisons and regulatory timeliness checks, ensuring that the final generated report text meets the standards for direct issuance.

[0020] Furthermore, the process of parsing the user-submitted engineering data, extracting multi-dimensional data items, and determining the target report type includes:

[0021] The format of the engineering data submitted by the user is decomposed and information is extracted to extract key engineering parameters. The key engineering parameters are reconstructed into an associated tuple structure containing parameter identifier name, value and physical unit to obtain the multidimensional data item.

[0022] By concatenating the parameter identifier names in the associated tuple structure, a feature string sequence is obtained;

[0023] The feature string sequence is mapped to a semantic feature vector and input into a fully connected classification network to calculate the probability distribution of the feature string sequence belonging to each preset report type;

[0024] The category corresponding to the highest probability value in the probability distribution is selected as the target report type.

[0025] Furthermore, based on the multidimensional data items, reference content is retrieved from the expert report knowledge base, including:

[0026] Generate a query text sequence based on the multidimensional data items, calculate the cosine similarity between the feature vector corresponding to the query text sequence and the feature vector in the vector index library of the expert report knowledge base, and obtain a semantic matching score;

[0027] The keyword matching score is obtained by statistically analyzing the word frequency status of the keywords contained in the query text sequence in the inverted index dictionary of the expert report knowledge base.

[0028] The semantic matching score and the keyword matching score are normalized and then fused to obtain a comprehensive evaluation value;

[0029] Candidate text fragments are recalled in descending order of the comprehensive evaluation value, and the candidate text fragments are used as the reference content.

[0030] In a preferred embodiment of the present invention, the shielding of item-specific factual information includes:

[0031] Determine the threshold range in which the comprehensive evaluation value falls, wherein the first threshold is greater than the second threshold;

[0032] When the comprehensive evaluation value is greater than or equal to the first judgment threshold, the project-specific factual information in the candidate text fragment is masked, while the chapter structure, indicator arrangement order, professional expression, data presentation format and argumentation depth are retained as reference benchmarks.

[0033] When the comprehensive evaluation value is less than the first judgment threshold and greater than or equal to the second judgment threshold, the specific parameter values ​​in the candidate text segment are stripped, and the peripheral syntactic structure and chapter organization method are retained as reference benchmarks.

[0034] When the comprehensive evaluation value is less than the second judgment threshold, the chapter title tag of the candidate text fragment is extracted as a reference benchmark;

[0035] The expert report knowledge base serves as both the training data source and the generation reference source for the vertical AI agent of the engineering consulting report. When used as a generation reference source, the expert report knowledge base does not directly provide the actual values ​​of the client name, project name, project location, site photos, monitoring data, equipment parameters, investment amount, and construction period for the current project.

[0036] Furthermore, the construction of the deterministic information closure, and the determination of whether the target entity object belongs to the deterministic information closure, if not, preventing the vertical AI agent in the engineering consulting report from generating passive value descriptions and inserting structured placeholders, includes:

[0037] Based on the entity rule file, extract the set of mandatory data frames corresponding to the target report type;

[0038] Based on the aforementioned list of entities that cannot be fabricated, a set of project fact entities that are prohibited from being freely generated by the model is extracted. The set of project fact entities includes at least the customer name, project name, specific geographical location of the project, coordinate information, photos of the project site, photos of the surrounding environment, production process flow chart, material balance sheet, equipment list, raw and auxiliary material list, land area, general layout plan, investment amount and construction period.

[0039] Calculate an entity state indicator function for the target entity object within the forced data framework set and the project fact entity set. The entity state indicator function is used to characterize whether the target entity object belongs to a deterministic information closure.

[0040] When the entity state indication function indicates that the target entity object does not belong to the deterministic information closure, it is determined that the target entity object does not belong to the deterministic information closure, wherein the deterministic information closure consists only of the following data:

[0041] (1) The multidimensional data item;

[0042] (2) Supplementary data confirmed by the user through the interactive interface;

[0043] (3) Fixed threshold data in the regulatory standards library;

[0044] (4) The result data calculated by the system based on the deterministic formula and the multidimensional data items;

[0045] Any numerical features and project-specific factual information in the reference content recalled from the expert report knowledge base are not incorporated into the deterministic information closure.

[0046] The vertical AI agent in the engineering consulting report is prohibited from generating value descriptions of the target entity objects. Instead, a structured placeholder containing a universally unique identifier, prompt description information, data type, unit requirements, and pending status is inserted at the corresponding position. The universally unique identifier, paragraph identifier, chapter identifier, and report identifier of the structured placeholder are written into the document structure.

[0047] In a preferred embodiment of the present invention, receiving supplementary user data includes:

[0048] Obtain a pre-constructed engineering data dependency graph, wherein the engineering data dependency graph uses the actual data entities corresponding to the multidimensional data items and the structured placeholders as nodes, and establishes directed edges with the dependency direction, dependency type, formula identifier, unit conversion rules and applicable conditions recorded by the pre-set domain prior association matrix. After loop detection, it is obtained and satisfies the directed acyclic constraint.

[0049] Extract the specific supplementary values ​​and universally unique identifiers from the user's supplementary data;

[0050] Using the universally unique identifier as the query key, the corresponding node is retrieved in the engineering data dependency graph. The retrieved node is marked as the original change node, and the specific supplementary value is written into the original change node.

[0051] Establish a replacement mapping relationship between the specific supplementary values ​​and the corresponding structured placeholders.

[0052] Further, the step of replacing the structured placeholders and generating the rewrite report text includes:

[0053] Starting from the original change node, perform a reachable node search along the directed edges in the engineering data dependency graph to extract the derived nodes that depend on the original change node, thus forming a set of derived nodes;

[0054] Perform a topological sort on the derived node set, and call the calculation formula carried by the directed edge in topological order to calculate the new value of each derived node;

[0055] The associated change area is determined based on the universally unique identifier, paragraph identifier, chapter identifier, and report identifier of the structured placeholder;

[0056] Replace the corresponding structured placeholders with the specific supplementary values, and merge the new values, calculation formulas, calculation basis, and the preceding context text of the associated change area into a local rewrite instruction prompt;

[0057] The local rewrite instruction prompt is input into the vertical AI agent of the engineering consulting report, a new text paragraph is output, and the associated change area is replaced in the first draft of the engineering report;

[0058] The partial rewriting is automatically triggered after the user submits supplementary data, and only the text areas associated with the structured placeholders and their derived nodes are rewritten. The partial rewriting does not trigger a full report regeneration; the content of chapters in the initial draft of the project report that are not associated with the structured placeholders remains unchanged.

[0059] In a preferred embodiment of the present invention, the rewritten report text is automatically reviewed, and modifications are made to issues that can be automatically corrected, while annotations and review comments are generated for issues that cannot be automatically corrected, including:

[0060] Scan the rewritten report text, extract numerical segments with the same physical meaning, and extract the indicator name, statistical object, spatial range, time scale, working condition stage, state before and after treatment, statistical caliber, unit and chapter source as caliber tags;

[0061] The numerical segments that meet the consistency comparison conditions of the caliber labels are selected and converted to the standard International System of Units according to the conversion benchmark to obtain the standard numerical sequence.

[0062] Calculate the range of the standard numerical sequence. When the range is greater than a preset tolerance threshold, it is determined that there is a data logic conflict, and the conflict information is recorded in the abnormal information cache area.

[0063] Scan the chapter numbers, figure numbers, format identifiers, legal and standard references and structured placeholders in the rewritten report text, identify format-related anomalies, legal timeliness-related anomalies, placeholder residue-related anomalies and anomalies requiring manual confirmation, and enter the anomaly information cache area;

[0064] The information in the anomaly information cache is structured and arranged to generate the review opinion list;

[0065] Based on the review opinion list, issues that can be automatically corrected and those that cannot be automatically corrected are distinguished; for errors in chapter numbering, errors in figure and table numbering, inconsistent formats, residual formatting of placeholder prompts, and regulatory standard references with identifiable alternative versions, automatic correction is performed; for regulatory application issues that lack factual data, conflicting values ​​of the same indicator, exceed reasonable thresholds and cannot determine the correct value, or require expert judgment, paragraph annotation information and modification opinions are generated;

[0066] Based on the type, quantity, severity level, and automatic correction status of the anomalies in the review comments list, the issue-ready status of the rewritten report text is determined as follows:

[0067] (a) When there are no standard alarms, data logic conflicts, placeholder residues, or manual confirmation anomalies in the review opinion list, the issue status is marked as issueable;

[0068] (b) When the list of review comments contains only automatically correctable anomalies and the automatic correction has been completed, the issue status is marked as issueable after correction;

[0069] (c) When any of the following issues are present in the review opinion list: specification out-of-bounds warning, data logic conflict, placeholder residue, or anomaly requiring manual confirmation, the issueability status is marked as unissueable or pending manual confirmation. The issueability review result is then generated.

[0070] Furthermore, the automated review also includes:

[0071] Obtain the set of boundary thresholds consisting of the normatively mandated thresholds and empirically reasonable thresholds;

[0072] A reasonable numerical range is determined based on the set of boundary thresholds, and it is determined whether the indicator calculation results in the rewritten report text fall within the reasonable numerical range.

[0073] When the calculated result of the indicator does not fall within the reasonable value range, a standard violation alarm or data anomaly risk warning is generated based on whether the corresponding threshold belongs to a mandatory standard threshold or an empirically reasonable threshold. The mandatory standard threshold is derived from mandatory limit clauses in current laws, standards, or technical specifications; the empirically reasonable threshold is derived from reference limit clauses in historical expert report statistics, industry experience rules, or internal enterprise review rules. The anomaly information is then entered into the anomaly information cache for use in generating the review opinion list and the issueability review result.

[0074] Furthermore, the automated review also includes:

[0075] Traverse the rewritten report text and extract the sequence of legal and standard names and document numbers;

[0076] In the regulatory standard timestamp database of the regulatory standard library, the obsolescence status identifier corresponding to the regulatory standard name is checked by matching according to the standard name, standard number, version year, alias field and substitution relationship field;

[0077] If the standard referenced in the text has been replaced by an updated standard or marked as obsolete, the name and document number of the latest version of the standard are retrieved as replacement information and entered into the exception information cache to generate the review opinion list and the issueability review result.

[0078] This invention also includes a method for generating reliable data in engineering consulting reports, characterized by comprising:

[0079] Based on the aforementioned multidimensional data items, supplementary data confirmed by the user through the interactive interface, fixed threshold data in the regulatory standard library, and result data calculated by the system based on deterministic formulas, a deterministic information closure is constructed.

[0080] The target entity object to be generated is compared with the deterministic information closure;

[0081] When the target entity object belongs to the deterministic information closure, the vertical AI agent in the engineering consulting report is allowed to generate a corresponding description;

[0082] When the target entity object does not belong to the deterministic information closure, the generation of value descriptions is prohibited, and structured placeholders are inserted.

[0083] Specifically, the item-specific factual information and numerical features in the reference content recalled from the expert report knowledge base are not included in the deterministic information closure.

[0084] This invention provides a method for generating AI agents and reviewing the issueability of engineering consulting reports. It has the following beneficial effects:

[0085] 1. This invention limits the deterministic information closure by introducing mandatory entity rules and a list of non-fabricable entities. During the report generation phase, the system intercepts the model's generation of project parameters lacking data sources and replaces them with structured placeholders. This mechanism physically decouples text statement generation from factual data assignment, preventing the model from falsifying underlying engineering data when generating professional reports and ensuring the authenticity of core factual indicators in the report.

[0086] 2. This invention utilizes a pre-constructed engineering data dependency graph to handle data association changes. Upon receiving supplementary data for structured placeholders, the system, based on the topological sorting rules of the directed acyclic graph, sequentially calls the associated calculation formulas to update the values ​​of all derived nodes and triggers a partial rewriting of the associated text. This processing method ensures that when engineering parameters are modified, all associated calculation results can be synchronously converted, eliminating data omissions and association logic errors that are easily generated during manual calculation and modification.

[0087] 3. This invention establishes an automated review mechanism covering numerical consistency, threshold boundaries, and regulatory timeliness. By extracting numerical segments for standard international unit conversion and range comparison, combined with verification of the obsolescence status of regulatory standard timestamp databases, the system can identify anomalies such as data logic contradictions and invalid reference standards in rewritten report texts. Coupled with automatic correction and annotation prompt strategies for different anomaly categories, this improves the efficiency of the generated report's regulatory review, ensuring it conforms to the actual issuance standards of the engineering consulting industry. Attached Figure Description

[0088] Figure 1 This is a schematic diagram of the method execution architecture of the present invention;

[0089] Figure 2 This is a schematic diagram of the overall process of the AI ​​intelligent agent generation and issueability review method for engineering consulting reports of the present invention;

[0090] Figure 3 This is a schematic diagram illustrating the process of parsing engineering data, determining report types, and retrieving entity rule files according to the present invention.

[0091] Figure 4 This is a schematic diagram of the dual-path retrieval and reference benchmark generation process of the present invention;

[0092] Figure 5 This is a schematic diagram of the structured placeholder insertion process of the present invention;

[0093] Figure 6 This is a schematic diagram of the supplementary data-driven cascade regeneration process of the present invention;

[0094] Figure 7This is a schematic diagram illustrating the process of automatically verifying the rewritten report text, generating a list of review comments, and generating the issueability review results of the present invention.

[0095] Figure 8 A diagram showing the comparison of the chief engineer's review time and the number of revisions before submission for different report preparation methods;

[0096] Figure 9 A diagram illustrating the comparison of average processing time for key stages of report preparation under different report preparation methods. Detailed Implementation

[0097] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0098] See attached document Figure 1 This invention provides a method for generating an AI agent and reviewing the issueability of engineering consulting reports. This method can be executed collaboratively by an engineering consulting report vertical AI agent, an interaction interface module, a multimodal parsing module, a dual-path retrieval module, a placeholder processing module, a cascaded regeneration module, a consistency review module, and a data-driven module. The above modules are used to illustrate the specific execution architecture of this method and do not limit this method to relying on a specific hardware form.

[0099] The vertical AI agent for engineering consulting reports is trained or fine-tuned based on historical engineering consulting reports that have been issued or reviewed by experts. It is used to learn the chapter organization, argumentation depth, indicator arrangement order, professional expression, data presentation format, and review rules of different types of engineering consulting reports. Unlike methods that directly call general-purpose large language models to generate report text, the vertical AI agent for engineering consulting reports is constrained by an expert report knowledge base, entity rule files, a list of non-fabricable entities, and deterministic information closure judgment rules during the report generation process.

[0100] The interactive interface module is used to receive engineering data, structured placeholder supplementary data and manual confirmation information submitted by external terminals, and output the generated, supplemented, reviewed and corrected report documents and the corresponding review comments list to external terminals.

[0101] The multimodal parsing module communicates with the interactive interface module to perform text recognition, image feature extraction, and table structure transformation on engineering data, and calculates the report type classification intent of the target engineering report based on the extraction results.

[0102] The dual-path retrieval module communicates with the data-driven module to perform vector semantic matching and keyword precision matching in the expert report knowledge base, respectively, to obtain a comprehensive evaluation value. Based on the comprehensive evaluation value, it retrieves business reference fragments, expression structure features, data presentation formats, or text structure frameworks. Historical project factual information in the expert report knowledge base is only used as a reference for structure, expression, and argumentation depth, and is not used as a source of values ​​for the current engineering project.

[0103] The placeholder processing module is used to compare the entity data extracted by the multimodal parsing module with the entity rule file and the list of unfabricated entities corresponding to the target report type during the report text generation process. When the comparison result shows that the current generation position involves missing entities or project fact entities that cannot be freely generated by the model, it prevents the vertical AI intelligent agent of the engineering consulting report from generating passive value descriptions and inserts a structured placeholder carrying a universally unique identifier at the corresponding text position.

[0104] The cascaded regeneration module is used to construct the engineering data dependency graph within the report text. After the interactive interface module receives the supplementary entity data for the structured placeholders, it locates the original change node in the engineering data dependency graph based on the universal unique identifier, determines the associated change area affected by the original change node, and performs local rewriting on the associated change area.

[0105] The consistency review module is used to perform entity value range verification, context value consistency verification, standard and specification timeliness verification, placeholder residue check, and format consistency check on the rewritten report text, and compile the abnormal information that fails the verification into a review opinion list; for issues that can be automatically corrected, automatic modification is performed; for issues that cannot be automatically corrected, paragraph annotation information and manual confirmation prompts are generated.

[0106] The data-driven module communicates with the vertical AI agent for engineering consulting reports, the multimodal parsing module, the dual-path retrieval module, the placeholder processing module, the cascaded regeneration module, and the consistency review module. It is used to build and maintain an expert report knowledge base, a regulatory and standard timestamp library, entity rule files corresponding to various reports, a list of non-fabricable entities, a domain prior association matrix, a formula rule base, and a set of boundary thresholds. The expert report knowledge base also serves as a training data source and a generation reference source for the vertical AI agent for engineering consulting reports.

[0107] See attached document Figure 2 The method for generating and reviewing the AI ​​intelligent agent for engineering consulting reports of the present invention includes the following steps.

[0108] S1 receives the initial input of engineering data, including text, drawings and tables, through the interactive interface module, calls the multimodal parsing module to parse the initial input of engineering data to extract multidimensional data items, maps the multidimensional data items to the corresponding target report type, and retrieves the entity rule file and non-fabricable entity list corresponding to the target report type from the data-driven module.

[0109] S2, based on the extracted multidimensional data items, trigger the dual-path retrieval module to calculate the semantic matching score and keyword matching score in the expert report knowledge base of the data-driven module respectively, and extract the reference paragraphs, expression structure features, data presentation format or basic directory structure that meet the conditions as reference benchmarks based on the comprehensive evaluation value and the set score threshold range.

[0110] Specifically, the project-specific factual information in the reference paragraphs is masked so that the expert report knowledge base only provides references for report structure, expression, and argumentation depth, and does not serve as a source of values ​​for the current engineering project.

[0111] S3, during the initial draft generation stage of the engineering report, continuously compares multidimensional data items, entity rule files, and the list of unfabricated entities using the placeholder processing module. When the generation instruction points to a target entity object that does not belong to the deterministic information closure, or points to a project fact entity in the list of unfabricated entities and the user has not provided corresponding data, it prevents the vertical AI agent in the engineering consulting report from generating passive value descriptions through decoding constraints, structured output verification, or post-processing replacement mechanisms, and inserts structured placeholders at the corresponding positions.

[0112] S4 receives supplementary data returned by the user for the structured placeholder through the interactive interface module, calls the cascade regeneration module to lock the original change node in the engineering data dependency graph according to the associated placeholder identifier carried by the supplementary data, extracts the affected derivative nodes along the dependency path, loads the supplementary data in the topological order, calculates the values ​​of the derivative nodes and rewrites the text area corresponding to the derivative nodes to obtain the rewrite report text.

[0113] The process of the user sending back supplementary data for the structured placeholder corresponds to the second round of user interaction. After the user submits the supplementary data, the cascading regeneration module automatically triggers local rewriting and limits the rewriting scope to the text area associated with the structured placeholder and its derived nodes.

[0114] S5 calls the consistency review module to automatically review the rewritten report text, extracts the numerical segments of the same indicator object in the entire document, calculates the numerical range and determines the consistency status; at the same time, it checks whether the indicator calculation results exceed the boundary threshold, checks the timeliness status of the cited laws and standards, and scans the residual status of chapter numbers, figure numbers, format identifiers and structured placeholders; it generates a review opinion list from the extracted abnormal information, automatically modifies issues that can be automatically corrected, generates paragraph annotation information and modification opinions for issues that cannot be automatically corrected, and outputs the issueability review result.

[0115] See attached document Figure 3 During step S1, the interactive interface module receives initial input of engineering data transmitted from an external terminal. This initial input includes unstructured text reports, project site photos, engineering drawings, bills of materials, and equipment lists. The multimodal parsing module parses this data to form the factual data foundation required for subsequent text generation and data verification.

[0116] S11, the multimodal parsing module performs format decomposition and information extraction on the received heterogeneous files. For plain text and image data, it calls the optical character recognition algorithm to extract the basic character matrix; for tabular data such as bills of materials and equipment details, it calls the table structure parsing operator to convert the physical row and column boundaries in the scanned document or spreadsheet into a logically nested data structure with corresponding table headers.

[0117] For engineering drawings, the multimodal parsing module extracts the drawing title, drawing number, scale, legend text, coordinate labels, equipment numbers, pipeline identifiers, process flow labels, and drawing column information, and converts the extracted results into drawing element data items. For spatial relationships, layout relationships, or process connection relationships that cannot be determined through image parsing, the multimodal parsing module does not generate definitive conclusions, but instead outputs the corresponding content as structured placeholders to be supplemented or confirmed.

[0118] For the specific implementation process of optical character recognition and the conversion of physical row and column boundaries of tables into logical nested data structures, those skilled in the art can use existing deep learning vision models and image processing operators. The principles and network training process are well-known technologies in this field and will not be elaborated here.

[0119] The multimodal parsing module further extracts key engineering parameters from the transformed basic character matrix and logically nested data structure, and reconstructs these key engineering parameters into an associative tuple structure containing parameter identifier names, numerical values, and physical units. In this implementation, the extracted multidimensional data items are represented as sets. ,in This represents the total number of extracted engineering features.

[0120] For each element in the set The extracted engineering parameter is identified by its name. For the corresponding specific numerical parameters, This refers to the dimensions and physical units associated with the numerical value. For parameters that only have a textual description and no dimensional attribute, The identifier is set to null to maintain the consistency of the data item structure. This tuple construction method maps user input to a computer-processable storage structure, providing a basis for comparison objects for subsequent data validation.

[0121] S12, the multimodal parsing module reads the sequence content of the extracted text features and performs classification intent calculation on the business requirements pointed to by the initial input of engineering data. The system is pre-configured to support an automatically generated engineering report type classification library, denoted as set. .

[0122] In one specific embodiment of the present invention, the multimodal parsing module first identifies the parameter identifiers in the extracted associated tuple structure. Perform string concatenation to form a sequence of characteristic strings. Feature string sequence It also includes the user's input in natural language, the uploaded file name, the file type identifier, the chapter title, the key paragraph text, the parameter units, and the business tags identified in the drawings or tables.

[0123] The multimodal parsing module is configured with a Transformer-based text encoder. This text encoder extracts the contextual relationships between engineering terms in the feature string sequence through a self-attention mechanism and maps them to a dimension of 1. High-dimensional semantic feature vector ,in The preset hidden layer dimension.

[0124] The multimodal parsing module will convert semantic feature vectors The input is a fully connected classification network. This fully connected classification network includes an input layer and a linear mapping layer, and calculates the probability distribution of the input content belonging to each preset report type using a normalized exponential function. The probability distribution calculation satisfies the formula:

[0125] ;

[0126] In the formula, Corresponding query input The probability determination result mapped to a specific report type, and satisfying the following conditions: And the sum of the probabilities of all categories in the classification library equals 1; The learning weight matrix parameters for the corresponding classification network have a dimension of . ; The corresponding classification network has a dimension of The bias parameter vector.

[0127] The system compares the maximum probability value in the probability distribution and determines the category corresponding to the maximum probability value as the target report type. When the maximum probability value is lower than the preset classification confidence threshold, the multimodal parsing module outputs multiple candidate report types and requests user confirmation through the interactive interface module. After user confirmation, the system uses the confirmation result as the target report type and continues with subsequent processing. This mechanism enables the system to classify the current task into specific business categories such as environmental impact assessment report, safety pre-assessment report, and feasibility study report based on project parameters, file type, and business context, even when the user has not explicitly specified the report type.

[0128] Furthermore, the aforementioned fully connected classification network requires model training before use. Training samples are derived from various types of previously issued and approved engineering reports. The system extracts historical feature string sequences from these reports and uses the actual business type of the sample report as the true label for supervised learning. During training, the model uses the cross-entropy loss function to calculate the predicted probability distribution. With real labels The network loss value between the two is calculated, and the text encoder parameters and weight matrix parameters are updated based on the backpropagation algorithm and gradient descent optimizer. and bias parameter vector This continues until the classification accuracy on the validation set reaches the preset convergence threshold.

[0129] S13, based on the target report type result obtained from intent classification calculation, the multimodal parsing module initiates an internal communication call instruction to the data-driven module. The underlying storage device of the data-driven module pre-maintains mandatory data framework sets corresponding to various types of engineering reports. In response to the call instruction, the data-driven module matches and downloads the entity rule file and the list of non-fabricable entities corresponding to the target report type.

[0130] The entity rule file corresponds to a list of information constraints for this report type, while the non-fabricable entity list corresponds to a list of project fact entities that are prohibited from being freely generated by the vertical AI agent of the engineering consulting report. These are stored as a preset set in the system. Elements inside the set These are the entity attributes required by industry experts when reviewing and approving the issuance of documents for this type of project. Each required entity object... It should include at least the entity name, entity alias set, chapter, data type, unit type, whether it is mandatory, placeholder prompt, validation rules, reasonable range, dependency identifier, legal source identifier, and manual confirmation identifier.

[0131] Through the aforementioned rule documents, the system transforms general compilation requirements into computable set comparison objects, ensuring that subsequent report generation is subject to clear data boundary constraints; in the event of missing data, the system can trigger generation interruption and placeholder insertion accordingly.

[0132] See attached document Figure 4 During step S2, the dual-path retrieval module queries the reference benchmark for the current project in the data-driven module. This process does not directly extract current project data from historical reports; instead, it utilizes the chapter structure, expression methods, and data presentation habits of previously issued expert reports to constrain the generation boundaries of the language model.

[0133] In this invention, the construction of the expert report knowledge base is interconnected with the training or fine-tuning of the vertical AI agent for engineering consulting reports. The data-driven module first acquires historical engineering consulting reports that have been issued or have passed expert review. These historical engineering consulting reports include at least one of the following: environmental impact assessment reports, safety pre-assessment reports, feasibility study reports, soil and water conservation plans, occupational disease hazard assessment reports, and emergency response plans.

[0134] The data-driven module performs chapter-level decomposition, paragraph content identification, data item extraction, legal and standard citation identification, chart number identification, and expert review conclusion identification on the engineering consulting history report. It then assigns report type tags, chapter function tags, industry category tags, argumentation depth tags, professional expression tags, data presentation tags, and review rule tags to the decomposed text fragments.

[0135] Based on historical engineering consulting reports with the aforementioned tags, the data-driven module constructs a training sample set and uses the training sample set to perform supervised fine-tuning, instruction fine-tuning, or retrieval-enhanced adaptation training on the basic language model. This enables the model to learn the chapter structure, argumentation path, usage of professional terms, order of indicators, data presentation format, and expert review focus of the engineering consulting report, thereby obtaining a vertical AI intelligent agent for engineering consulting reports.

[0136] The client name, project name, project location, monitoring data, equipment parameters, investment amount, construction period, and other project-specific factual information in the expert report knowledge base are not used as the source of factual values ​​for the current project. The project-specific factual information is masked, stripped, or downgraded to a structural reference after retrieval and recall to prevent historical project data from being mistakenly entered into the current project report.

[0137] S21, the data-driven module pre-decomposes and indexes the historically issued expert report set stored internally. Specifically, the data-driven module splits the historical report text into independent text fragments according to the physical directory structure, and retains the chapter title tags of each text fragment. The extracted set of all text fragments represents... ,in This represents the total number of text fragments in the knowledge base.

[0138] For this set of text fragments, the data-driven module is configured with a text embedding model. In a specific embodiment of the present invention, the text embedding model adopts a multi-layer bidirectional Transformer encoder architecture, including an embedding layer, a multi-layer self-attention mechanism layer, and an average pooling layer, for converting text sequences of unequal lengths into dense feature vectors of fixed dimensions.

[0139] Before constructing the vector library, this text embedding model requires fine-tuning training using engineering-specific corpora. Specifically, during training, the system constructs positive sample pairs from adjacent paragraphs in historical reports of the same category and negative sample pairs from randomly selected irrelevant paragraphs. After forward propagation into the network, a contrastive loss function is used to calculate the error between the predicted similarity and the sample label. The attention weight parameters within the Transformer encoder are then adjusted using a backpropagation algorithm to reduce the distance between semantically similar texts in the engineering domain within the vector space. After model convergence, the data-driven module calls the fine-tuned text embedding model to embed each text segment... Encoded as dense feature vectors To build a vector index library.

[0140] Meanwhile, the data-driven module calls the word segmentation tool to perform Chinese word segmentation on the text fragments, removes invalid prepositions, conjunctions and general function words from the stop word list, extracts terminology with engineering attributes and builds an inverted index dictionary.

[0141] S22, after the multimodal parsing module outputs the target report type and multidimensional data items, the dual-path retrieval module receives the query text sequence for the current task. The system adopts a hybrid retrieval mechanism that combines semantic similarity evaluation and literal term overlap evaluation in parallel to reduce the probability of missing key features in complex engineering statements when using a single-path retrieval method.

[0142] In practice, the dual-path retrieval module first performs word segmentation preprocessing on the query text sequence, filters out stop words, and then constructs a set of query keywords. Simultaneously, the text embedding model described above is invoked again to map the entire query text sequence into query feature vectors with dimensions consistent with the vector index library. .

[0143] The dual-path retrieval module initiates parallel feature matching operations within the data-driven module. In the first feature matching path, the dual-path retrieval module employs a vector cosine similarity calculation mechanism to verify the query feature vector. With each dense feature vector in the vector index library The inner product spatial distance is used to obtain the semantic matching score, which represents the semantic proximity of the context. The calculation logic satisfies the formula. .

[0144] In the second feature matching path, the dual-path retrieval module employs a sparse text retrieval mechanism, targeting the set of query keywords. The terminology elements in the dictionary are used to count the frequency of their corresponding words in the inverted index dictionary, resulting in a keyword matching score that characterizes the degree of hit rate for professional indicators. The expression for calculating this score satisfies:

[0145] ;

[0146] In the formula, Corresponding to specific keywords Inverse document frequency; Corresponding keywords In comparing text fragments The word frequency variable appears in the text; This variable represents the length of the number of characters contained in the currently compared text segment; This represents a statistical constant indicating the average length of text segments in the entire knowledge base. and This is a constant used to adjust the upper limit of matching frequency saturation and the length penalty effect, where The numerical range is usually set to take values ​​within the interval [1.2, 2.0]. It is usually set to a constant of 0.75.

[0147] S23, After the dual-path retrieval module obtains two independent matching results, it first scores the semantic matching results. Keyword matching score Normalize to the [0,1] interval to obtain the normalized semantic matching score. Normalized keyword matching score Then, a comprehensive evaluation value is obtained through linear fusion. The calculation constraints of the comprehensive evaluation value satisfy the formula:

[0148] ;

[0149] In the formula, This is a comprehensive evaluation value; The normalized semantic matching score; To convert the keyword matching score to the interval [0,1] using the maximum-minimum normalization algorithm; The weighting parameter, used to balance semantic bias and precision terminology bias in vectors, is set between 0 and 1. In this embodiment, It can be adaptively configured according to the user's needs for data accuracy or text polishing.

[0150] The dual-path retrieval module recalls candidate text fragments based on descending order of comprehensive evaluation values ​​and determines the structural reference level to be distributed to subsequent processes based on pre-defined parameter threshold ranges. The data-driven module pre-stores a first and a second judgment threshold, with the first judgment threshold set to be numerically greater than the second judgment threshold. In specific application scenarios, the first judgment threshold can be set to 0.8 to 0.85, and the second judgment threshold can be set to 0.5 to 0.6.

[0151] Furthermore, before distributing any candidate text fragment to the generation queue, the dual-path retrieval module first invokes the project-specific information masking subunit to mask the project name, company name, geographical location, contact person, approval number, monitoring data, equipment model, drawing number, project investment amount, construction period, and other factual data that cannot be transferred to the current project in the candidate text fragment. After masking, the system only retains the chapter structure, indicator arrangement order, professional expression method, data presentation format, and argumentation depth of the candidate text fragment, which is used to constrain the underlying language model of the vertical AI agent of the engineering consulting report to generate the current engineering project report text.

[0152] When the comprehensive evaluation value of the recalled candidate text fragments is greater than or equal to the first judgment threshold, the dual-path retrieval module inputs the chapter function, syntactic structure, indicator arrangement order and expression depth of the text fragment into the generation queue, which serves as a reference constraint for the underlying language model of the vertical AI intelligent body of the engineering consulting report to generate the current engineering project report text; project names, company names, geographical locations, specific values, monitoring data, approval numbers, drawing numbers and other project-specific factual information in historical text fragments are not directly included in the current report text.

[0153] When the comprehensive evaluation value is between the first and second judgment thresholds, the dual-path retrieval module performs local masking filtering on the candidate text fragments, strips out the specific parameter values ​​carried in them, retains the peripheral syntactic structure and chapter organization, and sends the filtered structural feature sequence to the subsequent generation process to trigger local placeholder logic.

[0154] When the comprehensive evaluation value is less than the second judgment threshold, the dual-path retrieval module determines that the retrieved corpus is not sufficiently relevant to the current project, discards the paragraph text of the candidate text segment, and only extracts its chapter title tags and sends them to the generation queue as the structure tree skeleton of the current new report's basic directory framework. This threshold grading mechanism is used to limit the underlying language model of the vertical AI agent in engineering consulting reports from generating passive content based on its own parameters when there is a lack of highly relevant historical material support.

[0155] See attached document Figure 5 When executing step S3, the placeholder processing module applies entity state constraints to the generation process of the underlying language model of the vertical AI agent in the engineering consulting report, so as to reduce the risk of generating passive value descriptions when the vertical AI agent in the engineering consulting report lacks real engineering data support.

[0156] S31, the placeholder processing module receives the set of actual engineering multi-dimensional data items extracted by the multimodal analysis module, and combines it with the supplementary data confirmed by the user, the fixed threshold data in the regulatory standard library, and the result data calculated by the system based on deterministic formulas to construct a deterministic information closure currently held by the system. The deterministic information closure includes the engineering data submitted by the user, the supplementary data confirmed by the user, the fixed threshold data in the regulatory standard library, and the result data calculated by the system based on deterministic formulas. Numerical features recalled from the expert report knowledge base are not directly incorporated into the deterministic information closure; they are only used as a reference for expression structure, reasonable range indication, or anomaly verification.

[0157] The placeholder processing module retrieves the entity rule file and the list of entities that cannot be fabricated from the data-driven module. The entity rule file is used to limit the entity attributes that must be present in the current report type, and the list of entities that cannot be fabricated is used to limit the project fact entities that cannot be freely generated by the vertical AI agent of the engineering consulting report. The list of entities that cannot be fabricated includes at least the client name, project name, specific geographical location of the project, coordinate information, photos of the project site, photos of the surrounding environment, production process flow chart, material balance sheet, equipment list, raw and auxiliary material list, land area, general layout plan, investment amount, and construction period.

[0158] The placeholder processing module extracts a mandatory data framework set for the current report type and a set of project fact entities that are prohibited from being freely generated by the model. Based on the idea of ​​set subtraction, for any target entity object within the above sets, the placeholder processing module defines and calculates its entity state indicator function. When the target entity object belongs to a deterministic information closure, it indicates that the entity object has objective data support; when the target entity object does not belong to a deterministic information closure, it indicates that the entity object is in a state of missing data.

[0159] When the target entity is in a state of missing data, the placeholder processing module prevents the vertical AI agent in the engineering consulting report from freely generating value descriptions, specific descriptions, or alternative speculations for the target entity, and inserts structured placeholders in the corresponding positions.

[0160] S32, the placeholder processing module inputs the instruction context obtained by the system and the reference frame retrieved and retained from the expert report knowledge base into the underlying large language model to perform forward inference. In this embodiment, the underlying large language model is constructed using a decoder architecture based on a self-attention mechanism, which uses an autoregressive prediction mechanism to output the word sequence one by one when outputting the text sequence of the initial draft report.

[0161] To enforce generated content constraints, the system pre-extracts state indication functions. The corresponding target entity name, entity alias set, chapter tag, data type and unit type are used to construct an entity matching trie and entity slot identification rules.

[0162] In one implementation, when the underlying language model is a locally deployed or open decoding interface language model, the placeholder processing module attaches a state monitoring operator during the decoding process to monitor the generated historical word sequence in real time; when the underlying language model is an external interface model, the placeholder processing module achieves interception processing equivalent to state monitoring through structured output constraints, function calls, generation result verification, and placeholder replacement mechanisms.

[0163] In implementations deployed locally or with an open decoding interface, the state monitoring operator compares the tail string of the current generation context with the entity matching trie during the probability distribution stage of generating the next lexical. For external interface models, after the model outputs structured results, the placeholder processing module verifies the content corresponding to missing entities according to entity slot recognition rules, and replaces the corresponding text with structured placeholders when a missing entity is matched.

[0164] When the comparison result matches the target entity object The system retrieves its corresponding status indicator function. At this point, the placeholder processing module prevents the language model from freely generating value descriptions corresponding to the target entity object and calls the placeholder generation function to return a structured placeholder at the current position. For locally deployed language models with controllable decoding processes, the placeholder processing module intercepts through candidate lexical constraints or decoding constraint interfaces; for external large language model interfaces, the placeholder processing module achieves equivalent interception through pre-constraint template generation, structured function calls, regular expression validation, failure retries, and post-processing replacement mechanisms. Regarding the implementation mechanism of autoregressive prediction based on self-attention mechanism and the internal probability matrix calculation of the decoding layer in the underlying large language model, those skilled in the art can use existing generative pre-trained network architectures for construction. The internal weight calculation rules of the network are well-known technologies in this field and will not be elaborated here.

[0165] S33, after severing the text output link of the large language model, the placeholder processing module inserts a structured placeholder at the breakpoint position of the current text sequence. Internally, the structured placeholder is instantiated as a data tuple containing multi-dimensional mapping attributes, denoted as... Its internal field structure is set as follows:

[0166] ;

[0167] in, This is a universally unique identifier dynamically generated by the system based on the global timestamp and the current paragraph identifier code, used to provide a positioning anchor point when supplementing input later; The enumeration values ​​for the type constraints of the data required to supplement the missing entity include text, numerical, image media, or geographic coordinates. This is a natural language prompt description information extracted from the entity rule file, used to prompt the user to supplement the corresponding parameters; The status enumeration field for this structured placeholder has values ​​including at least "to be supplemented", "supplemented but pending verification", "verification failed", "filled", "confirmed", and "abandoned". Its initial value is set to the "to be supplemented" status.

[0168] Furthermore, the structured placeholder also includes a unit requirement field, which is used to record the physical unit, unit type, or unit format that the target entity object needs to meet when supplementing data by the user.

[0169] When generating the initial draft of the engineering report, the placeholder processing module will use structured placeholders. The paragraph identifier (paragraph_id), chapter identifier (chapter_id), and report identifier (report_id) are written into the document structure. For Word format report files, the system saves the above mapping relationship through hidden bookmarks, content controls, custom XML nodes, or invisible tags; for HTML or structured text format report files, the system saves the above mapping relationship through tag attributes or metadata fields. Through this mapping relationship, after users supplement data, the cascading regeneration module can reliably locate the corresponding chapter and paragraph based on the unique identifier of the placeholder.

[0170] The placeholder processing module executes the above logic until the entire report's structure is traversed. The system integrates the text fragments with inserted structured placeholders with existing data fragments to form a draft engineering report file. The interaction interface module retrieves this draft engineering report file and outputs it to external operators. This execution mechanism transforms the continuous text generation process into a controlled text stream with mandatory interception nodes. Through deterministic state indication operations and structured placeholder replacement, it reduces the probability of the system generating unused data for core indicator parameters.

[0171] See attached document Figure 6 During step S4, the cascade regeneration module processes the entity data added by the user for the structured placeholders and performs local dynamic updates on the report text.

[0172] S41, after the initial draft of the engineering report is generated and before receiving supplementary data for structured placeholders, the cascaded regeneration module calls the pre-built domain prior association matrix within the data-driven module to pre-construct the engineering data dependency graph within the report text. This engineering data dependency graph is physically and logically defined as a directed acyclic graph. .

[0173] After completing the construction of nodes and directed edges, the cascade regeneration module performs loop detection on the engineering data dependency graph. When a circular dependency is detected, the system deletes low-priority back edges according to the preset dependency priority, or merges multiple nodes that form a circular dependency into a composite node, so that the final engineering data dependency graph used for topology traversal and local rewriting satisfies the acyclic constraint.

[0174] in, The node set in the graph represents the actual data entities extracted from the initial draft of the engineering report, as well as the structured placeholders attached to it. The set of directed edges between nodes represents the business deduction and dependency relationships between different engineering data indicators. Each directed edge includes at least the source node identifier, target node identifier, dependency type, calculation formula, applicable conditions, unit conversion rules, source identifier of legal or empirical rules, and confidence level. Before performing text rewriting, the cascaded regeneration module prioritizes performing deterministic calculations based on the calculation formulas and unit conversion rules carried by the directed edges, and passes the calculation results as strongly constrained inputs to the language model.

[0175] In this embodiment, the cascaded regeneration module will generate the node set The entity identifiers in the matrix are mapped to the domain prior association matrix for comparison and query. When two nodes are found to have a causal relationship or logical sequence relationship in the prior association matrix, the cascade regeneration module establishes a directed edge between the corresponding two nodes.

[0176] The domain prior association matrix uses the project entity identifier as the row and column index. Matrix elements record whether a dependency exists between two project entities, the direction of the dependency, the type of dependency, the formula identifier, the unit conversion rule, the applicable conditions, and the priority. The cascade regeneration module establishes directed edges based on the dependency direction in the matrix elements and calls the corresponding calculation formula based on the formula identifier.

[0177] In specific implementation scenarios, if the pollution generation coefficient rule constraint in the data-driven module limits the exhaust gas emission value to be calculated based on the raw material consumption value, the cascaded regeneration module establishes a directed edge pointing to the latter between the node corresponding to the raw material consumption and the node corresponding to the exhaust gas emission. For the memory allocation and graph structure storage representation method of the underlying data structure of the directed acyclic graph, those skilled in the art can apply existing adjacency matrix or adjacency list techniques. The basic data structure construction is a well-known technology in this field and will not be elaborated upon here.

[0178] S42, after the interactive interface module receives supplementary data for a specific structured placeholder sent by the external terminal, the cascaded regeneration module parses the supplementary data and extracts its accompanying universally unique identifier. And specific supplementary values. The cascaded regeneration module will use a universally unique identifier. As a query key, in the node set of the engineering data dependency graph The system performs a search, anchors the internal nodes of the graph corresponding to the universally unique identifier, and marks them as source change nodes. Source change nodes refer to engineering data-dependent graph nodes whose values ​​are directly triggered by user-added data.

[0179] After the cascaded regeneration module completes the parsing of the supplementary data, it first updates the internal state enumeration field of the node from the state to be supplemented to the state to be supplemented and awaiting verification. Subsequently, the cascaded regeneration module retrieves the type constraint enumeration value corresponding to the original change node. Perform data type consistency checks on the specific supplementary data.

[0180] When the supplementary data is numerical, the system further verifies the numerical format, physical units, unit conversion relationships, reasonable value ranges, and conflict status with existing data; when the supplementary data is text data, the system verifies null values, text length, and the integrity of sensitive fields; when the supplementary data is image media data, the system verifies the file format, resolution, file integrity, and its matching relationship with placeholder requirements; when the supplementary data is geographic coordinate data, the system verifies the latitude and longitude format, coordinate system identifier, and coordinate range.

[0181] When the verification fails, the cascade regeneration module intercepts the data writing, reports the verification failure exception to the external terminal, and updates the status enumeration field to the verification failure status. When the verification passes, the cascade regeneration module writes the specific supplementary value into the value storage field of the original change node and updates the internal status enumeration field of the node to the filled status.

[0182] S43, starting from the original change node, the cascaded regeneration module follows the directed edge set. The connection path is searched for reachable nodes, and all subsequent derived nodes that directly or indirectly depend on the original modified node data are extracted to form a set of derived nodes. Subsequently, the cascaded regeneration module generates the derived node set. Perform topological sorting and calculate the values ​​of derived nodes and rewrite the corresponding text regions in the topological order to ensure that the update results of upstream nodes are called before those of downstream nodes.

[0183] For sets For each derived node in the process, the cascaded regeneration module locates the physical paragraph position corresponding to the derived node in the initial draft of the project report file through the paragraph mapping identifier stored inside the node, and defines the physical paragraph as the associated change area to be modified.

[0184] S44, the cascade regeneration module first calls the deterministic calculation sub-unit set in the cascade regeneration module, or calls the formula rule library in the data-driven module, based on the specific supplementary values ​​entered in the original change node, the constraint derivation formulas corresponding to the derived nodes, the unit conversion rules and applicable conditions, to calculate the new values ​​of the derived nodes; then, it merges the new values, constraint derivation formulas, calculation basis and the preceding context text of the associated change area to construct the local rewrite instruction prompt.

[0185] Specifically, the partial rewrite instruction prompt contains a three-dimensional input structure: the first dimension is the source data input area for forced replacement, the second dimension is the original paragraph text retention area, and the third dimension is the generation logic constraint area extracted based on the domain prior association matrix. The cascaded regeneration module calls the underlying language model generation interface of the vertical AI agent of the engineering consulting report to input the partial rewrite instruction prompt into the vertical AI agent of the engineering consulting report. Under the constraint of the partial rewrite instruction prompt, the vertical AI agent of the engineering consulting report outputs a new text paragraph that incorporates the association change area with the newly supplemented values ​​based on an autoregressive mechanism. After obtaining the new text paragraph, the cascaded regeneration module performs text block replacement in the initial draft file of the engineering report.

[0186] This mechanism transforms single-point data supplementation into synchronous updates of logically related paragraphs, and limits the scope of text rewriting through paragraph mapping and dependency graphs, thereby reducing the computational cost of global regeneration and the content changes of unrelated paragraphs.

[0187] See attached document Figure 7 During step S5, the consistency review module performs automated verification on the rewritten report text to extract potential data conflicts, out-of-bounds values, and abnormal timeliness of specifications.

[0188] S51, the consistency review module scans the entire report text and uses a pre-built global data dictionary to extract sets of numerical segments distributed across different paragraphs for indicators with the same physical meaning. While extracting numerical segments, the consistency review module simultaneously extracts the indicator name, statistical object, spatial range, time scale, operating stage, pre- and post-treatment status, statistical caliber, unit, and chapter source as caliber labels; only numerical segments whose caliber labels meet the consistency comparison criteria are grouped into the same standard numerical sequence. .

[0189] To ensure the validity of subsequent numerical calculations, the consistency review module uses regular expressions to perform noise removal and cleaning on the extracted numerical fragments, stripping away non-numeric modifiers and removing thousands separators, and converting them into floating-point data. Let the original numerical sequence extracted and cleaned for a certain indicator object be... ,in This represents the total frequency of the indicator throughout the text. To eliminate dimensional differences caused by varying writing styles across different paragraphs, the consistency review module, based on the conversion benchmarks set in the data dictionary, uniformly converts elements carrying different physical units in the original numerical sequence to the standard International System of Units (SI), resulting in an aligned standard numerical sequence. .

[0190] The conformance review module calculates the range of the standard numerical sequence. and range value With respect to the preset tolerance threshold Perform a comparison. Tolerance threshold. The allowable fluctuation range of calculation precision is determined based on the data dictionary for different physical quantities. In this embodiment, for the basic measurement indicators, The value range is set to

[10] times the baseline order of magnitude of this indicator. -4 10 -3 Within the range. When the range value Greater than the tolerance threshold When the system determines that there is a data logic conflict in the description of the indicator object in the whole text, the consistency review module will enter the location of the conflicting paragraphs and the numerical difference into the abnormal information cache area.

[0191] In step S52, the conformance review module initiates a request to the data-driven module to retrieve the set of boundary thresholds for the current target report type. The boundary threshold set includes two categories: mandatory standard thresholds and empirically reasonable thresholds. Mandatory standard thresholds are derived from current regulations, standards, or technical specifications and are used to determine whether there are any mandatory breaches. Empirically reasonable thresholds are derived from historical expert reports, industry experience rules, or internal review rules and are used to indicate data anomaly risks. The boundary threshold set internally stores reasonable numerical ranges for specific process parameters or design indicators, denoted as... ,in and Representing the first The lower and upper limits of each indicator.

[0192] The consistency review module extracts the indicator calculation results from the text and determines whether the result falls within the corresponding reasonable value range. If the indicator calculation result is detected to be less than... or greater than If the corresponding threshold is a mandatory threshold stipulated by the standard, the consistency review module generates a standard out-of-bounds alarm message; if the corresponding threshold is an empirically reasonable threshold, the consistency review module generates a data anomaly risk warning message. The system writes the specific parameter names, threshold sources, deviation directions, and deviation amounts into the anomaly information cache. This mechanism is used to verify whether the generated data deviates from the basic boundaries of engineering physical reality or does not meet the constraints of current standards.

[0193] S53, the conformance review module uses a combination of regular expressions and named entity recognition models to traverse the report text field and extract the sequence of mentioned regulatory and standard names and document numbers.

[0194] In one specific embodiment of the present invention, the system is configured with a named entity recognition model based on the BERT-CRF architecture. This model includes a text embedding layer, a bidirectional Transformer encoding layer, and a conditional random field decoding layer.

[0195] During inference, the text embedding layer segments the input report text into character sequences and maps them to word vectors. The bidirectional Transformer encoding layer uses a multi-head self-attention mechanism to extract character-level bidirectional contextual semantic features and outputs the hidden layer state vector corresponding to each character. The Conditional Random Field decoding layer receives the hidden layer state vectors and, combined with the internally learned label transition probability matrix, outputs the BIO label sequence with the highest global probability. Here, B-LAW represents the first character of the regulation name, I-LAW represents characters within the regulation name, and O represents non-entity characters. The conformance review module extracts the complete regulatory standard name entity by merging consecutive B-LAW and I-LAW labels.

[0196] Furthermore, the aforementioned named entity recognition model undergoes fine-tuning training before use. When constructing the training sample set, text sequences containing real regulatory boundaries are manually labeled from past engineering report texts, serving as both sample input and real labels. During training, the negative log-likelihood of a conditional random field is used as the loss function to calculate the relationship between the model's predicted sequence and the true label. The error between the two layers is calculated, and the network weight parameters of the BERT layer and CRF layer are updated through the backpropagation algorithm until the entity recall rate of the model on the validation set tends to stabilize and converge.

[0197] After extracting the names of regulations and standards, the conformity review module initiates a timeliness verification query in the regulations and standards timestamp database maintained within the data-driven module. The regulations and standards timestamp database records the promulgation date, implementation date, and repeal status of each standard, and further records the standard name, standard number, version number, substitution relationship, replaced standard, implementation date, repeal date, current status, alias field, applicable report type, applicable industry, and data source. The data-driven module maintains the regulations and standards timestamp database through periodic updates or manual triggering, and generates a version snapshot after each update to ensure the traceability of regulations timeliness review results.

[0198] When querying the regulatory standard timestamp database, the consistency review module performs multi-field matching based on standard name, standard number, version year, alias, and substitution relationship fields. When the standard name and standard number are inconsistent, the standard number and version year are prioritized as the verification basis. Subsequently, the consistency review module compares the extracted regulatory standard name with the status identifier in the regulatory standard timestamp database.

[0199] If the query finds that the cited standard has been superseded by a newer standard or marked as obsolete, the consistency review module extracts the relevant text segment and retrieves the name and document number of the latest version of the standard from the regulatory standard timestamp database. If the query returns no match, it indicates that the cited standard is not included in the current database, and the consistency review module marks the cited standard as being in an unknown state. The consistency review module packages the expired citation information, recommended updated replacement information, and unknown state alerts and enters them into the exception information cache.

[0200] S54, the consistency review module performs structured organization on data logic conflicts, numerical out-of-bounds alarm information, and expired reference information collected in the anomaly information cache. The consistency review module categorizes and merges anomaly types and corresponding text paragraph positions, generating a final review comment list. Each record in the review comment list includes at least the anomaly number, anomaly type, severity level, chapter, paragraph position, original text excerpt, basis for the anomaly, suggested modifications, whether automatic correction is possible, automatic correction result, manual confirmation status, and associated placeholder identifier.

[0201] Furthermore, the consistency review module categorizes the scanned anomalies into format anomalies, regulatory timeliness anomalies, placeholder residue anomalies, data logic anomalies, numerical out-of-bounds anomalies, and anomalies requiring manual confirmation, and records the above anomalies in the anomaly information cache area.

[0202] Furthermore, the conformity review module categorizes abnormal information into those that can be automatically corrected and those that require manual confirmation. For issues such as incorrect chapter numbers, missing figure / table numbers, inconsistent unit formats, residual placeholder prompts, incorrect standard name formats, and definitive regulatory version replacements, the conformity review module invokes the automatic correction subunit to perform text corrections and records the content before correction, the content after correction, and the basis for correction in the review comments list.

[0203] For anomalies such as lack of factual data, conflicting values ​​of the same indicator, indicator calculation results exceeding reasonable thresholds and the correct value cannot be determined, contradictory measured data, missing design parameters, disputes over the applicability of specifications, or incomplete user information, the consistency review module does not directly modify the main text of the report, but generates annotation information in the corresponding paragraphs and includes the matter in the manual confirmation items.

[0204] Furthermore, the consistency review module determines the issuerability status of the rewritten report text based on the type, quantity, severity level, and automatic correction status of the anomalies in the review comment list. When there are no specification out-of-bounds warnings, data logic conflicts, placeholder residues, or anomalies requiring manual confirmation in the review comment list, the issuerability status is marked as issuerable.

[0205] When there are only automatically correctable anomalies and the automatic correction is completed, the issueable status is marked as corrected and issueable; when there are anomalies such as specification out-of-bounds alarms, contradictory measured data, missing design parameters, disputes over specification applicability, incomplete user information, or still requiring manual confirmation, the issueable status is marked as pending manual confirmation or unissueable, and the corresponding issueability review result is generated.

[0206] The interactive interface module obtains the list of review comments, the results of the issueability review, and the final report file generated by the system after automatic correction or annotation, and sends it to the external terminal via network communication protocol. This execution process assists the verification personnel in locating text anomalies and reduces the probability of missed detections during full manual review by using cross-dimensional comparison, standardized timeliness tracking, and traceable anomaly arrangement.

[0207] Specific application examples:

[0208] Taking the generation process of an environmental impact assessment report for a certain electroplating industrial park as an example, the user inputs through the interactive interface module ("Help me prepare a planning environmental impact assessment; the park is 2.5 square kilometers, with 200 lines, for zinc plating, chromium plating, and nickel plating"), and uploads the planning text, current status monitoring report, and relevant engineering data. After parsing, searching, masking, placeholder, supplementing, partially rewriting, and automatic review, the system outputs a partially updated planning environmental impact assessment report, a list of review opinions, and the results of the approval review.

[0209] This process verifies that the present invention can complete the generation of the initial draft report, interception of missing data, cascading partial rewriting after the second round of supplementary data, and automatic review and diversion without using historical project fact data in the expert report knowledge base as the source of values ​​for the current engineering project.

[0210] After receiving the user input and uploaded materials, the interactive interface module performs format decomposition and information extraction on the user input text, planning text, current status monitoring report, and related attachments. It extracts multi-dimensional data items such as park area, number of production lines, process type, planning text name, and monitoring report name, and determines the target report type as a planning environmental impact assessment report based on these multi-dimensional data items. Subsequently, the data-driven module retrieves the entity rule file and non-fabricable entity list corresponding to the planning environmental impact assessment report.

[0211] The dual-path retrieval module searches historical reports on environmental impact assessments of electroplating industrial park planning in the expert report knowledge base based on the multidimensional data items, calculates semantic matching scores and keyword matching scores respectively, and then integrates them according to the aforementioned comprehensive evaluation formula:

[0212] ;

[0213] in, This is a comprehensive evaluation value. To normalize the semantic matching score, To normalize the keyword matching score, For weight parameters. In one embodiment, the normalized semantic matching score is 0.86, the normalized keyword matching score is 0.78, and the weight parameter is... If the score is 0.6, the overall evaluation value is 0.6 × 0.86 + 0.4 × 0.78 = 0.828. Since this overall evaluation value is greater than the first judgment threshold of 0.8, the dual-path retrieval module recalls the corresponding candidate text fragments and masks historical project names, company names, geographical locations, monitoring data, approval numbers, and other project-specific factual information, retaining only the chapter structure, indicator arrangement order, professional expression style, data presentation format, and argumentation depth as a reference benchmark for generating the current engineering project report.

[0214] When generating the initial draft of the engineering report, the placeholder processing module determines the entity status based on the entity rule file and the list of non-fabricable entities. When it detects that target entity objects such as customer name, specific project location, project site photos, galvanizing cleaning wastewater generation, and chromic acid mist generation rate do not belong to deterministic information closures, the placeholder processing module prevents the vertical AI intelligent agent of the engineering consulting report from generating passive value descriptions and inserts structured placeholders in the corresponding paragraphs.

[0215] The structured placeholders include a universally unique identifier, a descriptive message, data type, unit requirements, and the status to be supplemented. For example, in the engineering analysis section, insert (Please supplement: galvanizing cleaning wastewater generation, unit...). Insert structured placeholders (please specify the project's location) in the Project Overview section and (please specify project site photos) in the Attachment List section.

[0216] After the user supplements the structured placeholders with data in the second round, the cascaded regeneration module locates the original change node in the engineering data dependency graph based on the universally unique identifier carried by the structured placeholders. In one embodiment, the user supplements the amount of galvanizing cleaning wastewater generated as 160. The cascaded regeneration module writes the data into the corresponding source change node and identifies the water balance table, wastewater generation calculation section, pollution source strength calculation section and treatment measures section that have dependencies on it as derived nodes.

[0217] Subsequently, the cascaded regeneration module extracts the set of derived nodes along the directed edges in the engineering data dependency graph, calls the corresponding calculation rules in the formula rule base according to the topological order, calculates the new values ​​of each derived node, and only performs partial rewriting on the paragraphs related to engineering analysis, water balance, wastewater generation calculation and treatment measures, without regenerating the entire chapters that are unrelated to the supplementary data.

[0218] The consistency review module automatically reviews the rewritten report text, extracting numerical fragments of the same indicator object from different chapters and uniformly converting them into a standard numerical sequence. In one embodiment, the system extracts the galvanizing cleaning wastewater generation of 160 from the engineering analysis chapter, water balance chapter, and pollution source strength calculation chapter, respectively. 160.2 and 160 Based on the aforementioned method for calculating the range, the range of the standard numerical sequence is 0.2. If the preset tolerance threshold is 0.5 If so, the consistency review module determines that the indicator does not constitute a data logic conflict.

[0219] If the same indicator is found to be recorded as 168 in a certain chapter during subsequent reviews... If the range of the standard numerical sequence is greater than the preset tolerance threshold, the consistency review module will record the anomaly in the anomaly information cache and generate corresponding modification opinions in the review opinion list.

[0220] Meanwhile, the consistency review module calls the boundary threshold set to determine whether the calculated index results fall within a reasonable value range, and checks the regulatory standards cited in the report in the regulatory standard timestamp library to see if they have been replaced or marked as obsolete. For errors in chapter numbering, missing figure / table numbers, inconsistent formats, placeholder prompts indicating formatting remnants, and regulatory standards cited with identifiable alternative versions, the consistency review module performs automatic corrections.

[0221] For regulatory application issues that still lack factual data, have conflicting values ​​for the same indicator, exceed reasonable thresholds and cannot be determined correctly, or require expert judgment, the consistency review module generates paragraph annotation information and modification suggestions, and forms a list of review comments. The interactive interface module ultimately outputs a report document that has undergone partial rewriting and automatic review, a list of review comments, and the results of the issueability review.

[0222] Furthermore, to verify the effectiveness of this invention in terms of report preparation efficiency and review workload, the same planning environmental impact assessment report task was used to compare the manual method, the general large language model-assisted method, and this method. The comparison indicators included the manual input time in the report generation stage, system processing time, number of revision rounds before approval, and chief engineer's review time. In the manual method, the preparation of an environmental impact assessment report takes approximately 8 hours, the chief engineer's review time is approximately 3 to 4 hours, and the number of revision rounds before approval is typically 3 to 5.

[0223] When using this method, the user's main input is the first round of submitting materials and the second round of supplementing placeholder data, which takes about 1 hour for the user, about 10 minutes for system processing, about 1 hour for the chief engineer's review, and about 1 to 2 rounds of revisions before approval. The above data illustrates that, under the same task scenario, this method reduces the workload of repetitive manual writing, full-text searching and modification, and the chief engineer's item-by-item verification through structured placeholders, cascading partial rewriting, and automated review processes.

[0224] See attached document Figure 8 The figure compares the processing time of manual methods, general large language model-assisted methods, and the proposed method in the main stages of report preparation. The horizontal axis of the figure represents the four stages of data parsing, draft generation, anomaly detection, and text rewriting, while the vertical axis represents the processing time. The legends correspond to the manual method, the general large language model-assisted method, and the proposed method, respectively.

[0225] From the appendix Figure 8 As can be seen, this method maintains a low processing time in the initial draft generation stage, and shortens the processing time in the anomaly detection and text rewriting stages through automatic review and partial rewriting mechanisms, avoiding changes to the content of unrelated paragraphs due to full text regeneration.

[0226] See attached document Figure 9 The figure shows a comparison of the results of manual review, general large language model-assisted review, and the proposed method in terms of chief engineer review time and number of revision rounds before approval. The bar chart represents the chief engineer review time, the line chart represents the number of revision rounds before approval, and the legend corresponds to the chief engineer review time and the number of revision rounds before approval, respectively.

[0227] From the appendix Figure 9 As can be seen, this method has completed placeholder residue checks, data consistency checks, reasonable range checks, and regulatory timeliness checks before the report is submitted for review. It can modify issues that can be automatically corrected in advance and form paragraph annotations and review opinion lists for issues that cannot be automatically corrected, thereby reducing the amount of repetitive troubleshooting work during the chief engineer's review stage.

[0228] Through the above specific application examples and comparative verifications, it can be seen that the present invention can achieve expert report knowledge base reference, historical project factual information masking, missing data placeholder interception, cascading partial rewriting after the second round of supplementary data, and automatic review and modification diversion of report text during the engineering consulting report generation process, which can improve data traceability, text consistency and review processing efficiency during the engineering consulting report generation process.

[0229] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.

Claims

1. A method for generating and reviewing the issueability of engineering consulting reports using AI-powered intelligent agents, characterized in that... include: Obtain a vertical AI agent for engineering consulting reports based on historical engineering consulting reports that have been issued or reviewed by experts, and build an expert report knowledge base. Parse the engineering data submitted by the user, extract multi-dimensional data items, determine the target report type, and retrieve the corresponding entity rule file and list of non-fabricable entities; Based on the multidimensional data items, reference content is retrieved from the expert report knowledge base, and item-specific factual information is masked. When generating the initial draft of the engineering report, a deterministic information closure is constructed for the target entity objects corresponding to the entity rule file and the list of non-fabricable entities. The deterministic information closure consists of the multidimensional data items, supplementary data confirmed by the user, fixed threshold data in the regulatory standard library, and result data calculated by the system based on deterministic formulas. Determine whether the target entity object belongs to the deterministic information closure. If not, prevent the generation of passive value description and insert a structured placeholder. Receive supplementary data from the user, replace the structured placeholders, and generate a rewrite report text; The rewritten report text is automatically reviewed, and the results of the approval review are output.

2. The method for generating and reviewing the issueability of engineering consulting reports using AI intelligent agents according to claim 1, characterized in that, The process involves parsing user-submitted project data, extracting multi-dimensional data items, and determining the target report type, including: The format of the engineering data submitted by the user is decomposed and information is extracted to extract key engineering parameters. The key engineering parameters are reconstructed into an associated tuple structure containing parameter identifier name, value and physical unit to obtain the multidimensional data item. By concatenating the parameter identifier names in the associated tuple structure, a feature string sequence is obtained; The feature string sequence is mapped to a semantic feature vector and input into a fully connected classification network to calculate the probability distribution of the feature string sequence belonging to each preset report type; The category corresponding to the highest probability value in the probability distribution is selected as the target report type.

3. The method for generating and reviewing the issueability of engineering consulting reports using AI intelligent agents according to claim 1, characterized in that, Based on the aforementioned multidimensional data items, reference content is retrieved from the expert report knowledge base, including: Generate a query text sequence based on the multidimensional data items, calculate the cosine similarity between the feature vector corresponding to the query text sequence and the feature vector in the vector index library of the expert report knowledge base, and obtain a semantic matching score; The keyword matching score is obtained by statistically analyzing the word frequency status of the keywords contained in the query text sequence in the inverted index dictionary of the expert report knowledge base. The semantic matching score and the keyword matching score are normalized and then fused to obtain a comprehensive evaluation value; Candidate text fragments are recalled in descending order of the comprehensive evaluation value, and the candidate text fragments are used as the reference content.

4. The method for generating and reviewing the issueability of engineering consulting reports using AI intelligent agents according to claim 3, characterized in that, The obscured item-specific factual information includes: Determine the threshold range in which the comprehensive evaluation value falls, wherein the first threshold is greater than the second threshold; When the comprehensive evaluation value is greater than or equal to the first judgment threshold, the project-specific factual information in the candidate text fragment is masked, while the chapter structure, indicator arrangement order, professional expression, data presentation format and argumentation depth are retained as reference benchmarks. When the comprehensive evaluation value is less than the first judgment threshold and greater than or equal to the second judgment threshold, the specific parameter values ​​in the candidate text segment are stripped, and the peripheral syntactic structure and chapter organization method are retained as reference benchmarks. When the comprehensive evaluation value is less than the second judgment threshold, the chapter title tag of the candidate text fragment is extracted as a reference benchmark; The expert report knowledge base serves as both the training data source and the generation reference source for the vertical AI agent of the engineering consulting report. When used as a generation reference source, the expert report knowledge base does not directly provide the actual values ​​of the client name, project name, project location, site photos, monitoring data, equipment parameters, investment amount, and construction period for the current project.

5. The method for generating and reviewing the issueability of engineering consulting reports using AI intelligent agents according to claim 1, characterized in that, The construction of the deterministic information closure, and the determination of whether the target entity object belongs to the deterministic information closure, if not, preventing the vertical AI agent in the engineering consulting report from generating passive value descriptions and inserting structured placeholders, includes: Based on the entity rule file, extract the set of mandatory data frames corresponding to the target report type; Based on the aforementioned list of entities that cannot be fabricated, a set of project fact entities that are prohibited from being freely generated by the model is extracted. The set of project fact entities includes at least the customer name, project name, specific geographical location of the project, coordinate information, photos of the project site, photos of the surrounding environment, production process flow chart, material balance sheet, equipment list, raw and auxiliary material list, land area, general layout plan, investment amount and construction period. Calculate an entity state indicator function for the target entity object within the forced data framework set and the project fact entity set. The entity state indicator function is used to characterize whether the target entity object belongs to a deterministic information closure. When the entity state indication function indicates that the target entity object does not belong to the deterministic information closure, it is determined that the target entity object does not belong to the deterministic information closure, wherein the deterministic information closure consists only of the following data: (1) The multidimensional data item; (2) Supplementary data confirmed by the user through the interactive interface; (3) Fixed threshold data in the regulatory standards library; (4) The result data calculated by the system based on the deterministic formula and the multidimensional data items; Any numerical features and project-specific factual information retrieved from the expert report knowledge base are not incorporated into the deterministic information closure. The vertical AI agent in the engineering consulting report is prohibited from freely generating value descriptions of the target entity object. Instead, a structured placeholder containing a universally unique identifier, prompt description information, data type, unit requirements, and pending status is inserted at the corresponding location. The universally unique identifier, paragraph identifier, chapter identifier, and report identifier of the structured placeholder are then written into the document structure.

6. The method for generating and reviewing the issueability of engineering consulting reports using AI intelligent agents according to claim 5, characterized in that, The receipt of supplementary data from the user includes: Obtain a pre-constructed engineering data dependency graph, wherein the engineering data dependency graph uses the actual data entities corresponding to the multidimensional data items and the structured placeholders as nodes, and establishes directed edges with the dependency direction, dependency type, formula identifier, unit conversion rules and applicable conditions recorded by the pre-set domain prior association matrix. After loop detection, it is obtained and satisfies the directed acyclic constraint. Extract the specific supplementary values ​​and universally unique identifiers from the user's supplementary data; Using the universally unique identifier as the query key, the corresponding node is retrieved in the engineering data dependency graph. The retrieved node is marked as the original change node, and the specific supplementary value is written into the original change node. Establish a replacement mapping relationship between the specific supplementary values ​​and the corresponding structured placeholders.

7. The method for generating and reviewing the issueability of engineering consulting reports using AI intelligent agents according to claim 6, characterized in that, The process of replacing the structured placeholders and generating the rewrite report text includes: Starting from the original change node, perform a reachable node search along the directed edges in the engineering data dependency graph to extract the derived nodes that depend on the original change node, thus forming a set of derived nodes; Perform a topological sort on the derived node set, and call the calculation formula carried by the directed edge in topological order to calculate the new value of each derived node; The associated change area is determined based on the universally unique identifier, paragraph identifier, chapter identifier, and report identifier of the structured placeholder; Replace the corresponding structured placeholders with the specific supplementary values, and merge the new values, calculation formulas, calculation basis, and the preceding context text of the associated change area into a local rewrite instruction prompt; The local rewrite instruction prompt is input into the vertical AI agent of the engineering consulting report, a new text paragraph is output, and the associated change area is replaced in the first draft of the engineering report; The partial rewriting is automatically triggered after the user submits supplementary data, and only the text areas associated with the structured placeholders and their derived nodes are rewritten. The partial rewriting does not trigger a full report regeneration; the content of chapters in the initial draft of the project report that are not associated with the structured placeholders remains unchanged.

8. The method for generating and reviewing the issueability of engineering consulting reports using AI intelligent agents according to claim 1, characterized in that, The rewritten report text is automatically reviewed; issues that can be automatically corrected are modified; and issues that cannot be automatically corrected are marked with annotations and review comments, including: Scan the rewritten report text, extract numerical segments with the same physical meaning, and extract the indicator name, statistical object, spatial range, time scale, working condition stage, state before and after treatment, statistical caliber, unit and chapter source as caliber tags; The numerical segments that meet the consistency comparison conditions of the caliber labels are selected and converted to the standard International System of Units according to the conversion benchmark to obtain the standard numerical sequence. Calculate the range of the standard numerical sequence. When the range is greater than a preset tolerance threshold, it is determined that there is a data logic conflict, and the conflict information is recorded in the abnormal information cache area. Scan the chapter numbers, figure numbers, format identifiers, legal and standard references and structured placeholders in the rewritten report text, identify format-related anomalies, legal timeliness-related anomalies, placeholder residue-related anomalies and anomalies requiring manual confirmation, and enter the anomaly information cache area; The information in the anomaly information cache is structured and arranged to generate the review opinion list; Based on the review opinion list, issues that can be automatically corrected and those that cannot be automatically corrected are distinguished; for errors in chapter numbering, errors in figure and table numbering, inconsistent formats, residual formatting of placeholder prompts, and regulatory standard references with identifiable alternative versions, automatic correction is performed; for regulatory application issues that lack factual data, conflicting values ​​of the same indicator, exceed reasonable thresholds and cannot determine the correct value, or require expert judgment, paragraph annotation information and modification opinions are generated; Based on the type, quantity, severity level, and automatic correction status of the anomalies in the review comments list, the issue-ready status of the rewritten report text is determined as follows: (a) When there are no standard alarms, data logic conflicts, placeholder residues, or manual confirmation anomalies in the review opinion list, the issue status is marked as issueable; (b) When the list of review comments contains only automatically correctable anomalies and the automatic correction has been completed, the issue status is marked as issueable after correction; (c) When any of the following issues are present in the review opinion list: specification out-of-bounds warning, data logic conflict, placeholder residue, or anomaly requiring manual confirmation, the issueability status is marked as unissueable or pending manual confirmation. The issueability review result is then generated.

9. The method for generating and reviewing the issueability of engineering consulting reports using AI intelligent agents according to claim 8, characterized in that, The automated review also includes: Obtain the set of boundary thresholds consisting of the normatively mandated thresholds and empirically reasonable thresholds; A reasonable numerical range is determined based on the set of boundary thresholds, and it is determined whether the indicator calculation results in the rewritten report text fall within the reasonable numerical range. When the calculated result of the indicator does not fall within the reasonable value range, a standard over-limit alarm or data anomaly risk warning is generated based on the corresponding threshold being a mandatory standard threshold or an empirically reasonable threshold. The mandatory standard threshold is derived from mandatory limit clauses in current laws, standards, or technical specifications; the empirically reasonable threshold is derived from reference limit clauses in historical expert report statistics, industry experience rules, or internal review rules of the enterprise, and is entered into the anomaly information cache area for generating the review opinion list and the issueability review result.

10. The method for generating and reviewing the issueability of engineering consulting reports using AI intelligent agents according to claim 8, characterized in that, The automated review also includes: Traverse the rewritten report text and extract the sequence of legal and standard names and document numbers; In the regulatory standard timestamp database of the regulatory standard library, the obsolescence status identifier corresponding to the regulatory standard name is checked by matching according to the standard name, standard number, version year, alias field and substitution relationship field; If the standard referenced in the text has been replaced by an updated standard or marked as obsolete, the name and document number of the latest version of the standard are retrieved as replacement information and entered into the exception information cache to generate the review opinion list and the issueability review result.

11. A method for generating reliable data in engineering consulting reports, characterized in that, include: Based on the aforementioned multidimensional data items, supplementary data confirmed by the user through the interactive interface, fixed threshold data in the regulatory standard library, and result data calculated by the system based on deterministic formulas, a deterministic information closure is constructed. The target entity object to be generated is compared with the deterministic information closure; When the target entity object belongs to the deterministic information closure, the vertical AI agent in the engineering consulting report is allowed to generate a corresponding description; When the target entity object does not belong to the deterministic information closure, the generation of value descriptions is prohibited, and structured placeholders are inserted. Specifically, the item-specific factual information and numerical features in the reference content recalled from the expert report knowledge base are not included in the deterministic information closure.

Citation Information

Patent Citations

  • Engineering technology standard intelligent identification and verification method based on multi-source information fusion

    CN121981085A

  • Text generation method based on large language model

    CN122113890A